Evidence strength, not automated truth.
Iraq OSINT continuously collects public reporting about Iraq in Arabic, English, and Hebrew, groups materially similar headlines, and calculates an explainable evidence score. The score is a triage aid. It is not a factual verdict, political rating, or substitute for human investigation.
1. Collection
The default deployment queries GDELT’s public news index in Arabic and English, filters GDELT for Hebrew-language sources, and reads Google News RSS searches for Iraq in Arabic, English, and Hebrew. Operators can add lawful public RSS or Atom feeds through the EXTRA_RSS_FEEDS_JSON environment variable. The deployment also runs a scheduled collection function. The main public feed keeps an exact rolling 48-hour window, the Public Officials Monitor keeps a separate exact rolling 24-hour window, and Policy & Research Watch retains Iraq-focused analysis from its curated research/publication registry for up to 90 days. The archive reduces gaps across serverless cold starts, but no deployment can guarantee complete continuous coverage.
The system does not bypass paywalls, impersonate users, access private groups, ingest hacked datasets, or collect content requiring authentication. It stores titles, descriptions supplied by feeds, publisher metadata, timestamps, links, and derived classifications—not complete copyrighted articles.
2. Hebrew-to-English translation
Hebrew records preserve the original Hebrew headline and source URL. The system attempts a machine translation into English and labels it as machine generated. The default package uses a public translation-memory service as a best-effort fallback; operators can configure Google Cloud Translation for a managed production service.
Translation can be unavailable, incomplete, delayed, or wrong. Names, quotations, negation, security terminology, and politically sensitive wording require review against the original Hebrew. When translation fails, the item remains in Hebrew and is marked as pending rather than receiving invented English text.
3. Normalization and deduplication
URLs are normalized by removing common tracking parameters. Exact duplicate URLs and normalized duplicate titles are removed. Records older than the exact rolling 48-hour main-feed window are discarded from the public dashboard. Confirmed Public Officials Monitor matches are retained separately for the latest 24 hours. Policy & Research Watch reports are stored separately for up to 90 days and do not widen the 48-hour live-news stream. A best-effort Netlify Blob archive preserves accepted records across midnight and serverless cold starts.
4. Story clustering
Headlines published within 48 hours are compared using normalized token overlap. Similar records are grouped into a story cluster. Public multi-domain support is based on the number of distinct publisher domains inside that cluster, not the number of copied articles. Domain diversity is not treated as proof of independent sourcing.
Arabic and Hebrew records can use bounded English machine translation of headlines for clustering while retaining the original title. Translation is best-effort and may be pending or imperfect, so cross-language clusters are candidate groupings rather than proof that reports are semantically identical or independently sourced. A human analyst should review important cross-language cases.
5. Topic and location tagging
Keyword dictionaries classify records into security, politics, economy, energy, humanitarian, climate, society, or other. Governorate names and common city aliases produce coarse location labels. These labels can be wrong and should be corrected by a reviewer.
The public pipeline is intentionally designed for governorate-level context. It should not expose exact coordinates of vulnerable people, active security positions, private residences, or information that could facilitate real-world targeting.
6. Evidence score
| Factor | Range | Meaning |
|---|---|---|
| Source provenance | 0–23 | Observable publisher identity, direct domain, and source type. An official source is strong evidence of what that institution announced, not independent confirmation of its claim. |
| Multi-domain support | 0–30 | Additional distinct publisher domains carrying materially similar reporting. This measures publisher-domain diversity only; it does not prove that the underlying sourcing is independent. |
| Recency | 0–10 | Recent records receive more value for live situational awareness. Age does not determine truth. |
| Attribution | 0–13 | Metadata that names a ministry, court, spokesperson, report, dataset, statement, or other identifiable basis. |
| Metadata completeness | 0–10 | Whether a useful public description is available rather than only a headline. |
| Direct publisher access | 0–5 | Whether the link resolves to the publisher rather than only an aggregator. |
| Caution penalties | 0 to −22 | Uncertainty language, sensational framing, or a casualty claim supported by only one domain. |
7. Labels
High evidence (75–100): stronger visible provenance, attribution, and/or multi-domain support. Details still require review.
Medium evidence (55–74): useful but incomplete support, often with limited independently verified sourcing.
Low evidence (0–54): weak metadata, a single source, uncertainty language, or caution penalties. Treat as an alert only.
Multi-domain: at least three distinct publisher domains in a materially similar cluster and a sufficient evidence score. This is not a claim that those publishers have independent underlying sources.
Unconfirmed: metadata contains allegation or uncertainty language. The label may be conservative or may miss subtler uncertainty.
8. Source grades
Source grades A–D are baseline provenance categories used by the calculation. They are not rankings of ideology or permanent judgments of accuracy. Publisher profiles must be reviewed and documented by the operator. Unrecognized domains receive a conservative “unrated” baseline.
9. Known limitations
- Indexes and RSS feeds do not cover every publisher or every story.
- Automated language, translation, topic, location, and cluster classifications can be wrong.
- Several outlets may repeat one original claim and appear independent when they are not.
- Metadata can omit corrections later added to the article.
- Headline similarity cannot reliably detect contradiction, satire, manipulated media, or coordinated influence activity.
- The system does not perform image forensics, geolocation, chronolocation, ownership analysis, or source-interview verification.
10. Editorial verification
Before converting a collected item into a published finding, an analyst should open and archive the source, identify the original claimant, seek independent primary evidence, check timestamps and context, compare Arabic, English, and Hebrew coverage, verify machine translations against the Arabic or Hebrew original, examine corrections, record uncertainty, and conduct a harm review.
11. Corrections and appeals
Every material correction should preserve the original publication time, state what changed, explain why, and update the score where relevant. Publishers and affected people should have a clear channel to challenge a classification without being required to disclose unnecessary personal information.
← Return to the live desk