Back to GDELT Cloud
Last updated July 5, 2026

Methodology

GDELT Cloud structures and enriches public event data so APIs, dashboards, and agents can use it without re-deriving the pipeline. This page explains where the data comes from, how often it updates, how we process it, and where it should not be the sole source.

Source universe

GDELT Cloud's primary raw inputs are the public GDELT 2.0 datasets: Events, the Global Knowledge Graph (GKG), and the Global Entity Graph (GEG). These are sourced from Google BigQuery, which hosts the official GDELT public datasets.

The hourly pipeline can also ingest native-source metadata from selected Arabic and Chinese RSS/news sitemap sources. These rows retain source language, country/region, provenance, and a per-source content policy, stored separately from the GDELT raw tables. That policy is honored, not overridden: sources marked blocked are fetched and logged but never admitted; metadata-only sources contribute discovery metadata (headline, URL, language) without full-article-body extraction; and only sources explicitly cleared for extraction have their article bodies processed.

Energy infrastructure and heavy-industry data are sourced directly from Global Energy Monitor's published registries — power and fuel assets (coal plants, renewable capacity, pipelines, LNG terminals, nuclear, and related categories) alongside GEM's Heavy Industry trackers (iron & steel, cement, chemicals, and iron ore). We obtain GEM data offline, separately from the GDELT ingest pipeline, and resolve asset owners through GEM's ownership graph into our unified entity registry so an entity's page shows the plants and mines it owns.

Corporate filings data (admin preview) is sourced from SEC EDGAR — the public-domain US securities filing system. We ingest the public filing index and structured XBRL financials hourly, within the SEC's published fair-access limits. EDGAR content is public domain; we cite the SEC as source. Corporate relationships (subsidiaries, suppliers, customers, jurisdictions) surfaced from filing text are derived through our proprietary AI pipeline and resolved into our entity registry — these derived relations are GDELT Cloud's interpretation, not statements by the filer.

Macro-economic time series (admin preview) are official U.S. economic statistics drawn from the Federal Reserve (FRED / ALFRED) and the originating agencies that publish them — the Bureau of Labor Statistics, Bureau of Economic Analysis, Census Bureau, and Treasury, among others. Observations are stored point-in-time (vintaged), so an as-of query returns the value as it was known on that date and never looks ahead. A minority of series carry third-party copyright (e.g. S&P, ICE, Dow Jones); we identify these and never store, serve, or use them, and macro text is never added to any embedding or fine-tuning corpus. This product uses data from the Federal Reserve Bank of St. Louis (FRED) but is not endorsed or certified by it.

Maritime vessel-flow signals are derived from public and commercial AIS vessel-tracking sources and expose only derived measures — chokepoint transit counts, dwell time, AIS-dark gaps, and last-known vessel positions and identity — across 11 maritime chokepoints (Hormuz, Bab-el-Mandeb, Malacca, Suez, Panama, Bosphorus, Gibraltar, Dover, Kerch, Taiwan, and the Danish Straits). We do not redistribute a raw AIS position feed or full track history. Maritime signals accrue from launch (June 2026) forward and are not backfilled; queries for dates before launch return no maritime coverage. Coverage is terrestrial-AIS — vessels are observed when within range of shore receivers, so open-ocean segments (which would require satellite AIS) are not captured, and coverage density varies materially by chokepoint: receiver coverage in the Persian Gulf is currently sparse, so Strait-of-Hormuz transits are effectively not observed today. Vessels are matched to Global Energy Monitor's LNG-carrier registry. The ports reference behind the port-disruption and proximity queries is the U.S. National Geospatial-Intelligence Agency's World Port Index (Pub 150) — a public-domain U.S. Government publication of roughly 3,800 ports with coordinates, harbor characteristics, and terminal facilities — with UN/LOCODEs from the UNECE code list.

AI-compute data (admin preview) exposes Epoch AI's published research datasets — notable AI models, ML hardware, AI data centers, chip-sales estimates, and AI-company metrics (funding, revenue, staff, and compute spend). These are loaded offline, separately from the GDELT ingest pipeline, and Epoch's organizations are resolved into our unified entity registry so an entity's page can show the models it developed, the hardware it makes, the data centers it owns, and its cumulative chip sales. Epoch AI data is published under CC-BY 4.0 and attributed to Epoch AI; it is a data source, not a subprocessor.

GDELT Cloud does not directly observe world events. We structure and enrich what reporters, public registries, and the GDELT Project's automated systems have already published.

Update frequency

The main ingest runs hourly on the top of the hour. Supporting jobs reconcile entity links, cluster membership, and GEG entity coverage on the same cadence, with a daily and weekly entity reconciliation pass that widens the window.

Native-source metadata is admitted under a capped incremental latency budget so it does not slow the main hourly cycle.

End-to-end latency, measured rather than estimated: over the seven days ending 2026-07-21, the median gap between when an article first became available to us — its GDELT feed timestamp for GDELT-discovered articles, or the publisher's own timestamp for direct native sources — and the coded event being available was 38 minutes, with a 90th percentile of 61 minutes. 10.7% of events were available within 15 minutes and 89.1% within an hour. The measurement covers about three-quarters of coded events; the rest are cross-cluster deduplication attachments whose originating article falls outside the window, and are excluded rather than guessed at. The floor is set by the hourly ingest cycle, not by processing time — an article that appears just after the top of the hour waits for the next run, so the median lands near half a cycle plus the pipeline's own duration. A long tail exists: events recovered by the hourly catch-up pass and by range backfills can land hours or days later, which is why we quote the full distribution rather than a single figure. The measurement harness and its exclusions will be released with our benchmark program. Variable upstream lag in GEG entity coverage can extend this for some entity-linked fields. Real-time freshness for each pipeline stage is published on the data status page.

Public entity, event, and story pages are cached for speed: data for the current UTC day refreshes on the hourly cadence, while pages for prior-day events and stories — which are settled — may be served from cache for up to 24 hours and refresh when reconciliation revises them. The REST API and MCP reflect the same hourly ingest cadence.

Atlas GPR — geopolitical-risk index (experimental)

Atlas GPR is a Geopolitical Risk index in the tradition of the Caldara–Iacoviello GPR, computed from GDELT Cloud's own data and normalized so that 100 is each place's own ~90-day norm — not a global average and not a fixed scale. A reading of 130 means a place is running about 1.3× its own typical level of geopolitical tension; each place is measured against itself.

It is published in two lenses that share the same deviation framework. The Events lens measures the share of a place's coded-event coverage that is tension-bearing — coercive intent before violence (threats, demands, force posture, sanctions) and violence already carried out (battles, explosions and remote violence, violence against civilians), with a severity split into Material and Verbal. This is our event-grounded, multilingual, drill-to-incident measure. The Attention lens is deliberately event-independent: the share of news attention — clustered narratives weighted by how many articles cover them — that a war/threat vocabulary identifies as tension-bearing. That is the same construction as the original newspaper-based GPR, applied to our multilingual narrative stream; for a single country it uses the narratives that name the country, exactly as the Fed's country-specific GPRC is built.

Both lenses roll up world → continent → region → country and move through time on a day-by-day basis. Below a per-level coverage floor a reading is withheld — returned as null with an insufficient-coverage flag — rather than guessed, so a thinly-covered geography never shows a fabricated zero. Baselines use a provisional ~90-day window; the index is labeled provisional and is not presented as an official long-run series. Atlas classifies nothing new on its own — it aggregates signals already produced upstream (coded events, or clustered narratives) into a share, normalizes to the baseline, and bands the result.

Atlas is experimental. The /api/v2/intelligence endpoints are gated on the Intelligence plan — they are no longer admin-only — but response shape and coverage may still change, and the country layer in particular is thin: measured over 107 days, only 2 countries clear the per-country coverage floor on 90 or more days and 12 clear it on most days. Treat world and continent as usable and country as preview.

Two things about that thinness are worth stating rather than burying. First, the coverage floor is applied to a single day, so a country producing a handful of events a day never clears it even though a week of its events comfortably would — a limitation of the daily series, not a judgement that the place is quiet. Second, the day-to-day corpus is not flat across the week: a Sunday carries far less news than a Wednesday, so a single-day share is noisier on weekends and the top bands fire disproportionately there.

A second construction, served in parallel rather than as a replacement, addresses both directly. It measures a place's share of the entire world's coded events over a fourteen-day window, normalized to that place's own frozen normal — which is what the Caldara–Iacoviello GPR and the Fed's country-specific GPRC actually do, and what GDELT's own documentation recommends. Dividing by corpus size rather than by the place's own coverage matters: a per-place denominator is topic volume, and it can make the index fall while the underlying events rise. Fourteen days is two whole weeks, so every reading contains each weekday exactly twice and the weekday cycle cancels by construction. There is no severity weighting, because it does not survive testing — the Fed's own intensity-weighted variant correlates 0.97 with plain counts, and fatalities are recorded on only a minority of qualifying events, so weighting by them would silently flatten half the index.

That construction carries a mandatory floor on the evidence behind its reference distribution, not just on the day being read. Below roughly a hundred qualifying events in the base window the index is not merely noisy, it is wrong: a country with a single event behind its normal reads eleven times normal forever. Above the floor the readings behave — the median is 91 and none fall outside half to double. So coverage is published as precision alongside every reading, and where the evidence is short the reading is withheld rather than reported. The deviation baseline is frozen over a fixed window rather than recomputed on a rolling one, because a rolling reference re-anchors a country to its own recent past — a country at war for the whole window acquires a wartime normal and reads unremarkable at the peak of what the index exists to detect.

Because the index is a ratio against a measured normal, the instrument that produced the events is part of the measurement. When the event coder changes generation, the distribution moves — across our country series the two generations disagree by more than 25% on roughly two thirds of them — so baselines are frozen per coder generation and every reading is normalized against its own. Each response says which. A series spanning a coder change is not one continuous series, and we label the break rather than splicing across it.

A band is a promise about how often a reading is unusual — roughly a tenth of a place's days above the ninth decile — and that promise requires a reference distribution with spread in it. Where a place's frozen cut-points are collapsed or unreachable, the band is withheld and the response says which case it is, rather than printing a label the distribution cannot support. The reading itself is unaffected. We would rather return a number without a word attached than a word that is wrong about its own meaning.

Atlas Posture — country condition index

Atlas Posture answers a different question from Atlas GPR, and the difference is the whole point. GPR asks how hot a place is running against its own norm — a deviation. Posture asks what state a place is actually in, scored against other countries and against fixed real-world reference points. A place can be at its own historical norm and still be in a severe condition; measuring only deviation would report a long war as unremarkable precisely because it has gone on long enough to become that country's normal. So Posture never normalizes a country against its own past. Own-history comparison appears in exactly one field: the trend arrow.

The score is a 0–100 composite of four pillars, published with fixed weights: Conflict & Security (0.35), Systemic Gravity (0.25), Markets & Economy (0.25) and Flows & Contagion (0.15). Each pillar is the mean of its sub-indicators, and each sub-indicator is a share or a per-event mass computed over a window — 30 days by default — by summing numerators and denominators across the window and dividing once, never by averaging daily figures. Sub-indicators are scored by where they fall in the pooled distribution across countries, with one deliberate exception: fatality intensity is scored against absolute anchors on the scale the conflict-research field recognizes (25 deaths per year, 1,000 per year), so a war reads as a war regardless of how the rest of the world is doing that quarter. Because our fatality counts come from our own conflict-event coding, they cover a broader set of deaths than battle-related-death datasets do; the anchors are reference points, not a claim of equivalence.

Every pillar is also computed twice more, split by whether an event's actors are entirely inside the country or span a border, giving an internal and an external reading beside the headline. Bands are fixed and published — Calm below 20, Steady 20–40, Elevated 40–60, High 60–80, Critical 80 and above — and each response also carries a neutral L1–L5 form of the same band, the pillar that is driving the score, the weights, the cut-points and the coverage floor in force.

Posture withholds more than it publishes, on purpose. A geography is scored only when its window clears a coverage floor, and each sub-indicator additionally requires a minimum denominator of its own: a share computed over two events is noise that can land anywhere in the distribution, and suppressing it matters more than filling the map. Roughly 60 countries currently clear the floor over 30 days; the rest return null with an insufficient-coverage flag. The one thing Posture will report as zero is a real measured zero — a well-covered country with no violence in the window is genuinely calm, and that is a finding rather than a gap.

A fifth pillar, Narrative & Information, is specified but deliberately not shipped: the inputs it needs are not available at the density a daily country panel requires, and publishing it on what does exist would manufacture a number. Responses name it as deferred rather than quietly omitting it. Like the rest of Atlas, Posture is provisional — its cross-country reference distribution is pooled over a short window because only a few months of consistently coded history exist.

Beside the event-derived pillars, Posture carries a second, slower axis: structural condition. Where the pillars above move on the news, this axis describes the standing institutional and material condition a country brings to any given week — governance quality and rule of law (World Bank Worldwide Governance Indicators), how democratic the regime is (V-Dem electoral, liberal, civil-liberties and free-expression indices), development, economy, and militarization. Governance and democracy are kept as separate dimensions on purpose: the World Bank measures whether a state is effective, V-Dem measures whether it is democratic, and a state can score well on one while scoring badly on the other. These inputs are annual and lagged, so this axis is deliberately slow — it is the condition a shock lands on, not the shock.

The structural axis is scored against fixed reference anchors rather than against the live field of other countries. This is a correctness requirement, not a preference: a cross-country rank is zero-sum, so it pins any world or regional aggregate near the middle by construction and lets a country's score improve only because another country deteriorated. Anchoring each indicator to fixed good and bad reference points on its own scale means a country's structural score depends only on its own condition, an aggregate for the world or a region becomes a real number that can move, and the axis stays commensurable with the rest of Atlas. Indicators whose meaningful range is multiplicative rather than additive, such as GDP per capita, are anchored on a log scale.

Event construction

Within each UTC ingest day, articles are grouped into stories and coded into structured events through a proprietary pipeline that combines AI classification, entity resolution, and our CAMEO+/ACLED taxonomies.

Coding spans both ACLED-style conflict events and CAMEO+ structured events across ten domains (political, crime, economic, corporate, technology, infrastructure, environment, health, demographic, and information). The taxonomies, output fields, and the event-metric formulas are all documented and published. The clustering, entity-resolution and fusion logic that assembles them is proprietary.

Taxonomy and scoring

GDELT Cloud uses two event taxonomies: ACLED-aligned conflict event types (political violence, protests, riots, strategic developments) and a CAMEO+ taxonomy that extends CAMEO across ten domain categories.

Scoring fields include the Goldstein scale (-10 to +10, populated for political and conflict events), magnitude (0-10), and the systemic_importance, propagation_potential, and market_sensitivity scores. Quad class and event-root codes are inherited from the underlying CAMEO/CAMEO+ structure.

Each of the four CAMEO+ metrics is produced by a fixed, published formula from sub-factors the coding model reads off the source text, recording a reason for each — so any value can be re-derived by a third party. magnitude routes by domain (conflict = 2 + 2·log10(deaths), following Richardson’s log-fatality convention; hazard, verbal and economic use anchored 0-10 tiers) and returns null, never zero, when no observable is found. systemic_importance is the equal-weighted mean of size, interconnectedness and irreplaceability, matching Basel BCBS G-SIB and ECB O-SII practice. propagation_potential averages five observed transmission conditions behind a channel gate, following the ERCS barrier model. market_sensitivity applies the reasonable-investor materiality test.

These are rubric-based scores structured by established frameworks, not measurements. A hard registry attribute — the observable that would make systemic_importance empirical — resolves for about 1 event in 45, and a traded instrument for about 1 in 67. So treat all four as ordinal ranking signals: not probabilities, not predicted price moves, and not scientific scales. Repeat coding of the same event moves magnitude by about 0.36 on its 0-10 scale and the three 0-1 metrics by about 0.03, so differences below roughly 0.72 and 0.07 on a single event are inside the noise. magnitude is comparable within a domain only; significance ranks events across domains. We adapt these frameworks; none of the institutions named endorses or reviews this work.

Story scope and clustering

Each story (cluster) carries a scope: local, national, regional, or international, derived from the geographic and source distribution of articles in the cluster.

Clustering operates on a UTC ingest-day window. Articles published on different days do not merge into the same cluster today, even when they describe the same ongoing incident. Cross-day continuity is on the roadmap and is flagged in the known limitations section below.

Risk context: screening lists, ownership exposure, China-Abroad

Plans that include Risk Context surface three reference layers alongside an entity's news, events, and assets. Screening-list context: entries from public restricted-party and sanctions lists — the U.S. Consolidated Screening List (which aggregates OFAC, BIS, and State Department lists), the UN Security Council and UK sanctions lists, and the U.S. DoD Section 1260H list — refreshed via scheduled diffs and stored bitemporally, so matches can be evaluated as of a given date. Ownership exposure: derived lenses (sanctions-linked, state-owned, China-linked) computed over the ownership relationships in our entity registry; exposure can be direct or flow through intermediate owners, and confidence varies with the depth and completeness of the underlying ownership data. China-Abroad: overseas development-finance projects and financier relationships from AidData's Global Chinese Development Finance dataset, resolved into our entity registry.

Risk Context is reference context, deliberately not sold or positioned as a compliance product. Matches are name/identifier-based against our entity resolution, list coverage is limited to the public lists above, and ownership chains are incomplete where public registries are thin. Every signal links to its underlying list entry or source document so it can be independently verified.

Risk Context output is NOT a compliance determination, NOT a system of record, and NOT legal advice. It must not be used as the sole basis for sanctions screening, restricted-party decisions, onboarding/KYC decisions, or any legal or regulatory obligation. Use it to prioritize what to investigate in an authoritative screening system, not to replace one.

Known limitations

Historical coverage is strongest from March 2026 forward. Earlier history is being backfilled but is incomplete; queries that span pre-March 2026 dates may return uneven coverage.

Source bias is real. GDELT's source universe over-indexes English-language and globally indexed outlets relative to local-language reporting, and some regions are systematically under-covered. GDELT Cloud inherits this bias.

Classification and deduplication are imperfect. CAMEO+ and ACLED coding rely on automated systems and may misclassify edge cases. Cross-cluster duplicates are flagged via a duplicate field but are not always perfectly resolved.

Cluster labels are generated by language models and consistency is not guaranteed across days or topics.

The story_count field is a cardinality of distinct stories within the queried window; it is not a sum across dates. Aggregating story_count across multiple dates without deduplication will overcount.

Appropriate use

GDELT Cloud is built for monitoring, triage, alerting, research acceleration, signal discovery, country and sector trend analysis, and agent workflows that benefit from structured event data with clean schema, entity links, and category coverage.

Not appropriate as the sole source for

Legal determinations, sanctions screening or restricted-party compliance decisions (Risk Context is reference context only — see the Risk Context section above), emergency-response decisions, sole-source investment decisions, individual-level risk decisions, or any claim requiring evidentiary certainty.

Treat GDELT Cloud output as analytical signal, not ground truth. High-impact conclusions should be corroborated with primary sources, official statements, or trusted reporting.

Schema and classifier changes

We may update schemas, classifiers, scoring logic, deduplication, clustering, and enrichment methods over time. Material changes that affect integrators are documented in the changelog. Where feasible we preserve backward-compatible fields and provide migration notes.

Methodology questions?

For deeper questions about data lineage, schema changes, or appropriate use, email us.