Evidence strength, not automated truth.
Iraq OSINT continuously collects public reporting about Iraq in Arabic, English, and Hebrew, groups materially similar headlines, and calculates an explainable evidence score. The score is a triage aid. It is not a factual verdict, political rating, or substitute for human investigation.
1. Collection
The default deployment queries GDELT’s public news index in Arabic and English, filters GDELT for Hebrew-language sources, and reads Google News RSS searches for Iraq in Arabic, English, and Hebrew. Operators can add lawful public RSS or Atom feeds through the EXTRA_RSS_FEEDS_JSON environment variable. The deployed collector runs whenever the feed is requested and while open browsers poll for updates. Operators may add a lawful external schedule if unattended collection is required. The two-day archive reduces gaps across serverless cold starts, but no deployment can guarantee complete continuous coverage.
The system does not bypass paywalls, impersonate users, access private groups, ingest hacked datasets, or collect content requiring authentication. It stores titles, descriptions supplied by feeds, publisher metadata, timestamps, links, and derived classifications—not complete copyrighted articles.
2. Hebrew-to-English translation
Hebrew records preserve the original Hebrew headline and source URL. The system attempts a machine translation into English and labels it as machine generated. The default package uses a public translation-memory service as a best-effort fallback; operators can configure Google Cloud Translation for a managed production service.
Translation can be unavailable, incomplete, delayed, or wrong. Names, quotations, negation, security terminology, and politically sensitive wording require review against the original Hebrew. When translation fails, the item remains in Hebrew and is marked as pending rather than receiving invented English text.
3. Normalization and deduplication
URLs are normalized by removing common tracking parameters. Exact duplicate URLs and normalized duplicate titles are removed. Records outside the two most recent Baghdad calendar dates are discarded from the public dashboard. A best-effort Netlify Blob archive preserves accepted records across midnight and serverless cold starts.
4. Story clustering
Headlines published within 48 hours are compared using normalized token overlap. Similar records are grouped into a story cluster. Corroboration is based on the number of distinct publisher domains inside that cluster, not the number of copied articles.
Hebrew records use the available English machine translation for clustering and classification while retaining the original title. Arabic-to-English translation is not included, so some Arabic and English reports about the same event may remain in separate clusters. A human analyst should merge important cross-language cases during editorial review.
5. Topic and location tagging
Keyword dictionaries classify records into security, politics, economy, energy, humanitarian, climate, society, or other. Governorate names and common city aliases produce coarse location labels. These labels can be wrong and should be corrected by a reviewer.
The public pipeline is intentionally designed for governorate-level context. It should not expose exact coordinates of vulnerable people, active security positions, private residences, or information that could facilitate real-world targeting.
6. Evidence score
| Factor | Range | Meaning |
|---|---|---|
| Source provenance | 0–23 | Observable publisher identity, direct domain, and source type. An official source is strong evidence of what that institution announced, not independent confirmation of its claim. |
| Independent-domain corroboration | 0–30 | Additional distinct domains carrying materially similar reporting. Syndicated copies on one domain do not count as independent confirmation. |
| Recency | 0–10 | Recent records receive more value for live situational awareness. Age does not determine truth. |
| Attribution | 0–13 | Metadata that names a ministry, court, spokesperson, report, dataset, statement, or other identifiable basis. |
| Metadata completeness | 0–10 | Whether a useful public description is available rather than only a headline. |
| Direct publisher access | 0–5 | Whether the link resolves to the publisher rather than only an aggregator. |
| Caution penalties | 0 to −22 | Uncertainty language, sensational framing, or a casualty claim supported by only one domain. |
7. Labels
High evidence (75–100): stronger visible provenance, attribution, and/or multi-domain support. Details still require review.
Medium evidence (55–74): useful but incomplete support, often with limited independent corroboration.
Low evidence (0–54): weak metadata, a single source, uncertainty language, or caution penalties. Treat as an alert only.
Corroborated: at least three distinct domains in a materially similar cluster and a sufficient evidence score. This does not prove that all domains are independent at the ownership or sourcing level.
Unconfirmed: metadata contains allegation or uncertainty language. The label may be conservative or may miss subtler uncertainty.
8. Source grades
Source grades A–D are baseline provenance categories used by the calculation. They are not rankings of ideology or permanent judgments of accuracy. Publisher profiles must be reviewed and documented by the operator. Unrecognized domains receive a conservative “unrated” baseline.
9. Known limitations
- Indexes and RSS feeds do not cover every publisher or every story.
- Automated language, translation, topic, location, and cluster classifications can be wrong.
- Several outlets may repeat one original claim and appear independent when they are not.
- Metadata can omit corrections later added to the article.
- Headline similarity cannot reliably detect contradiction, satire, manipulated media, or coordinated influence activity.
- The system does not perform image forensics, geolocation, chronolocation, ownership analysis, or source-interview verification.
10. Editorial verification
Before converting a collected item into a published finding, an analyst should open and archive the source, identify the original claimant, seek independent primary evidence, check timestamps and context, compare Arabic, English, and Hebrew coverage, verify machine translations against the Hebrew original, examine corrections, record uncertainty, and conduct a harm review.
11. Corrections and appeals
Every material correction should preserve the original publication time, state what changed, explain why, and update the score where relevant. Publishers and affected people should have a clear channel to challenge a classification without being required to disclose unnecessary personal information.
← Return to the live desk