Transparency

Methodology

How the historical dataset behind Terrorism Portal is assembled, deduplicated and classified.

1. Sources

The dataset combines three openly reusable sources, each covering a different slice of the historical record:

2. Assembly and deduplication

Records from the three sources are merged into one JSONL file (events_portal_v1.jsonl) with a common schema. Duplicate incidents — the same event described in more than one source — are identified using a combination of date, location, actor and identifier matching and merged into a single record, preserving the originating source where possible. Missing values are left unknown rather than inferred or invented.

3. Classification tiers

Every record carries a classification field, set at ingestion time based on which source flagged it and how directly:

Full breakdown of how many records fall into each tier is shown live on the Source & Confidence Criteria page and in the About section of the homepage — computed directly from the dataset, not hand-maintained.

4. Geocoding

Coordinates are taken from the source record when available. Roughly 3–4% of records have no valid coordinates (older or less-documented incidents especially). Those records remain in the dataset and the timeline, but are excluded from the map — see the map's own caption for the current count.

5. What "terrorism" means here

Terrorism is a contested category with no single legal or scholarly definition. Some records — political assassinations, insurgent attacks, one-sided violence in civil conflicts — sit near the boundary of competing definitions. Rather than adjudicate that boundary ourselves, the dataset preserves each source's own judgment as the classification field, and the interface surfaces that judgment instead of hiding it behind a single "confirmed" label.

6. Limitations