Methodology
How the historical dataset behind Terrorism Portal is assembled, deduplicated and classified.
1. Sources
The dataset combines three openly reusable sources, each covering a different slice of the historical record:
- Wikidata (CC0) — the historical backbone: structured records with identifiers, dates and locations going back to 1881.
- Wikipedia annual "List of terrorist incidents" pages (CC BY-SA) — additional terrorism-specific coverage not captured in Wikidata's structured fields.
- UCDP Georeferenced Event Dataset (CC BY, Uppsala Conflict Data Program) — organized-violence records, included only after conservative filtering (see §3).
2. Assembly and deduplication
Records from the three sources are merged into one JSONL file (events_portal_v1.jsonl) with a common schema. Duplicate incidents — the same event described in more than one source — are identified using a combination of date, location, actor and identifier matching and merged into a single record, preserving the originating source where possible. Missing values are left unknown rather than inferred or invented.
3. Classification tiers
Every record carries a classification field, set at ingestion time based on which source flagged it and how directly:
- Documented (
listed_terrorism) — the record appears on a dedicated "terrorist incidents" list (Wikipedia's annual lists, or an equivalent structured Wikidata list). - Related (
wikidata_terrorism_related) — Wikidata links the item to terrorism-related categories or properties, without it being on a dedicated incident list. - Unverified candidate (
terrorism_candidate) — drawn from UCDP's organized-violence data and flagged as a plausible terrorism candidate by keyword/pattern matching on the event description. These records have not been individually confirmed as terrorism and should be read as leads, not conclusions. - Unclassified — no classification field was set. Treated the same as "unverified" in the interface: not presented as documented terrorism.
Full breakdown of how many records fall into each tier is shown live on the Source & Confidence Criteria page and in the About section of the homepage — computed directly from the dataset, not hand-maintained.
4. Geocoding
Coordinates are taken from the source record when available. Roughly 3–4% of records have no valid coordinates (older or less-documented incidents especially). Those records remain in the dataset and the timeline, but are excluded from the map — see the map's own caption for the current count.
5. What "terrorism" means here
Terrorism is a contested category with no single legal or scholarly definition. Some records — political assassinations, insurgent attacks, one-sided violence in civil conflicts — sit near the boundary of competing definitions. Rather than adjudicate that boundary ourselves, the dataset preserves each source's own judgment as the classification field, and the interface surfaces that judgment instead of hiding it behind a single "confirmed" label.
6. Limitations
- Coverage is uneven across periods and regions — recent decades and English-language-documented regions are overrepresented relative to older or less-covered conflicts.
- Casualty figures can differ between sources; where sources disagree the dataset keeps the value from the record that was retained during deduplication rather than averaging or reconciling.
- The dataset is a snapshot assembled from public sources at specific points in time, not a continuously verified live registry outside of the (currently disabled) live evidence monitor.
See also: Editorial Policy · Corrections Policy · Source & Confidence Criteria