Indo-Pacific Record

Official defense and security texts, preserved as published and analyzed in context.

Collection is current; analysis is behind it. Records dated after 2026-08-24 are stored and awaiting screening. What these dates mean.

As published.

Corpus Guide

What the China Desk corpus holds, what it cannot tell you, what each field means, and how to cite it. Snapshot — 2026-08-26.

What this corpus holds

The China Desk corpus is a stored record of 3,574 items collected from official and authoritative Chinese defense and security publishers, held as they were captured. Each item keeps its original-language title, the address it came from, the date its source stated, and the collection metadata that describes how it arrived.

It is a record of what these institutions published. It is not a measure of what they did, and it is not a census of everything they published. An item is here because collection reached it and stored it, which is a narrower claim than it may look. Coverage is selective by design and incomplete in practice, and the sections below say where.

Each record has its own page carrying every stored field described in the dictionary. Records are reachable by publication week without JavaScript, and through search and filters on the the record.

Scope and limits

Four limits shape what any count from this corpus can support.

The corpus is concentrated in one institution

CMC Political Work Department accounts for 3,448 of 3,574 records — 96.5%. The remainder is spread across 2 other institutions and 4 outlets in total. Any comparison drawn across institutions in this corpus is a comparison of very unequal samples, and a pattern found here is a pattern in one publisher's output before it is anything else.

Repeats across outlets are stored once

De-duplication across sources is first-writer-wins: when more than one configured outlet carries the same item, it is stored once, under whichever source reached it first. A per-outlet total therefore counts the records an outlet contributed to this corpus, not how often that outlet published. Source counts are not publication-volume counts, and they should not be read as market share, prominence, or emphasis.

Some captures stored no source text

45 records hold an empty body capture. The original title, the source, the date and the original URL are all recorded for each; only the body is missing. An empty capture is a defect in what was stored, not evidence that the source page carried nothing, and consulting the original URL remains the way to establish what was published.

Most records carry no English rendering

2,239 of 3,574 records hold no machine-translated title. The same records — the same identifiers, not merely the same total — also hold no machine summary and no analysis-model record, and they are exactly the records outside the Analyzed state.

English text is a property of the analyzed subset, not of the corpus, and a search over English titles reaches only that subset. The original title is present on every record, and searching it reaches all of them.

How dates and collection work

Two different clocks run through this corpus, and confusing them produces wrong readings.

Publication dates are stated by the source. The date on a record is the date its publisher gave it. Nothing verifies that date against another authority, and it is collection-bounded: it can only fall inside the window collection actually reached. Records exist here from 2026-05-07 to 2026-08-26 because that is when collection ran, not because those dates bound anything the institutions did.

Collection timestamps are the pipeline's own clock, in UTC. They record when an item was stored, not when it was published and not when anyone read it. Because publication dates are stated on the source's calendar and collection timestamps are UTC, the two do not align, and a record is routinely stored on a later calendar date than the one it carries.

A date with no recorded run is not a quiet day. The run record shows which UTC dates carried at least one pipeline run. Where a date carries none, nothing was collected — which says nothing at all about what sources published. A gap in the run record is a gap in observation and must never be read as an observed absence.

The run record is thinner than it looks. A run row proves a pipeline run happened. It does not preserve which publication dates were sought, and it does not preserve historical per-source outcomes. So a run identifier cannot establish which sources were reached on a given day, and this corpus cannot support per-source reliability history. Current-run results and the vocabulary used to describe them are on Coverage.

Processing states

Every record is in exactly one of four processing states, and the four sum to the snapshot total. This matters more than it may appear: reading the corpus as though records were simply analyzed or not merges records that were screened and deliberately set aside with records that were never screened at all. Those are different claims about what is known.

The four processing states in this snapshot. Every stored record is in exactly one; the four sum to 3,574.
State Records What it means
Analyzed 1,335 records Passed relevance screening and completed analysis. English title and summary are machine-generated.
Not selected for analysis 1,402 records Screened and not selected for analysis. The original record is stored; no translation or summary was produced.
Awaiting screening 834 records Stored but not yet screened for relevance. No judgment of any kind has been made about this record.
Analysis incomplete 3 records Passed relevance screening, but analysis did not complete. The original record remains stored; no completed analysis is claimed.

A state describes how far a record traveled through the pipeline. It is not a judgment of the item's importance, and it carries no editorial meaning.

Machine-generated layers

Three things in this corpus were produced by software rather than captured from a source, and each is marked as such wherever it appears.

English titles and summaries are unreviewed machine output. Where a record carries an English title, a model produced it from the original title. Where it carries a summary, a model produced that from the stored source text. They are two separate artifacts, not one. No human has checked either against the original, and both are a reading of what was published rather than a verification of it. The original-language title is the authoritative title in every case.

Machine assessment is a triage cue, and only inside one state. Records that completed analysis carry a software flag marking them for closer review. It is meaningful only within the analyzed set, and it is displayed nowhere else. The stored value defaults to negative, so a negative reading cannot distinguish a record that was assessed and not flagged from one that was never assessed at all. For that reason a negative value is never shown outside the analyzed set, and it must never be read as evidence that an unscreened record was assessed and cleared.

Significant and Routine are editorial, and belong to editions, not records. Those two words label a PLA Watch edition — the editor's own classification of a week — and are applied by a person. They are not a machine output, they do not describe any corpus record, and they never appear on one. The pipeline's article-level flag and the editor's edition label sit at different rungs of the same ladder and are deliberately kept apart. The editions themselves are listed under Analysis.

Data dictionary

Every field a reader encounters on a record page, what it means, and where it stops being reliable. Stored column names appear beside the reader-facing label where knowing them helps.

17 fields. “Stored” means the value is held in the database as captured or as generated; “Derived” means it is computed when this page is built.
Field Origin What it means When it may be absent Main limitation
Record ID id Stored What it means. The identifier of a stored record inside this snapshot. When absent. Never absent. Limitation. A locator inside this snapshot, not a permanent public identifier. Identifiers are not contiguous — the highest is 3580 across 3,574 records — so a range of identifiers never describes the corpus.
Source-stated publication date published_date Stored What it means. The publication date as the source itself stated it. When absent. Never absent. Limitation. Source-stated and collection-bounded. It records what the source said, and it can only fall inside the window collection actually reached. Nothing verifies it against another authority.
Source outlet sources.display_name Stored What it means. The outlet that published the item, taken from the configured source that collected it. When absent. Never absent. Limitation. De-duplication across sources is first-writer-wins, so an item carried by several outlets is held once, under whichever source stored it first. Outlet totals are counts of stored records, not of publication volume.
Publishing institution institutions.display_name Stored What it means. The institution behind the outlet. When absent. Never absent in this snapshot. A source configured without an institution would leave it empty. Limitation. The corpus is heavily concentrated: CMC Political Work Department accounts for 3,448 of 3,574 records (96.5%). Institution totals describe this collection, not the wider field.
Original language sources.language_tag Stored What it means. The language of the original item. It is an attribute of the source, inherited by every record collected from it. When absent. Never absent. Limitation. Inherited from the configured source and never detected on the record itself. A source that published in a second language would still report its configured one.
Original title title_original Stored What it means. The item's title as published, stored as captured. When absent. Never absent. Limitation. Held as captured and never edited, translated in place, or normalized. It is the authoritative title for this record.
Stored source text text_original Stored What it means. Body text captured from the source page. When absent. Empty in 45 records. Limitation. Extraction can omit material or pull in unrelated page furniture, so no capture is a facsimile. An empty capture is a stored defect, not evidence that the page carried nothing.
Machine-translated English title title_english Stored What it means. A model's English rendering of the original title. When absent. Absent for 2,239 records. Limitation. Unreviewed machine output. No human has checked it against the original title.
Machine summary summary_english Stored What it means. A model's summary of the stored source text. When absent. Absent for 2,239 records. Limitation. Unreviewed machine output, and a reading of what was published rather than a verification of it.
Processing state passed_relevance, analyzed_at Derived What it means. How far a record traveled through screening and analysis. Every record is in exactly one of the four states above. When absent. Never absent; it is computed for every record. Limitation. Processing has four states, not two. Reading the corpus as analyzed-or-not merges records that were screened and set aside with records never screened at all.
Machine assessment Stored What it means. A software triage cue marking a record for closer review. When absent. Shown only for records in the Analyzed state. Limitation. Meaningful only inside the Analyzed state. The stored column defaults to a negative value, so a negative reading cannot tell an assessed record from one that was never assessed — and must never be read as evidence that an unscreened record was assessed and cleared.
Analysis model model_id Stored What it means. The model that produced the English title and summary. When absent. Absent for 2,239 records. Limitation. Names the model only. It does not record the model's configuration or how it behaved on this record.
Prompt version prompt_version Stored What it means. The analysis prompt in force when the record was analyzed. When absent. Absent for 253 of the 1,335 analyzed records. Limitation. Where it is absent, the exact prompt behind that record's English output cannot be established afterward.
Capture fingerprint content_hash Stored What it means. A hash taken once, at capture, over the stored original title and text. When absent. Never absent. Limitation. A capture-time fingerprint that is never recomputed. It records what arrived; it is not a continuing integrity guarantee, and a later correction to the stored text would leave it stale.
Collection run scrape_run_id Stored What it means. The pipeline run that stored the record. When absent. Never absent. Limitation. Run history preserves no target publication date and no historical source-level outcome, so a run identifier does not establish which sources were reached or which dates were sought.
Collection timestamp scraped_at Stored What it means. When the pipeline stored the record. When absent. Never absent. Limitation. A pipeline collection timestamp in UTC — not a reader access date and not a publication time. Publication dates are stated on the source's own calendar, so the two do not align.
Original URL url Stored What it means. The address the item was collected from. When absent. Never absent. Limitation. Recorded as collected. Whether it still resolves is not checked when this page is built.

Identifiers and citation guidance

Each record has an identifier and a page inside this snapshot. Both are locators, and neither is a permanent public name.

A record identifier locates an item within the 2026-08-26 snapshot. It is not a DOI, not an accession number, and not a public permalink; no resolver, registry, or persistence commitment stands behind it. Identifiers are not contiguous either — the highest is 3580 across 3,574 records — so a range of identifiers never describes this corpus, and no citation should quote one.

Because a record identifier is snapshot-scoped, a citation has to carry the snapshot with it. Citing a record means citing two things: the source text as its publisher issued it, and the record as this corpus holds it. Those are separate claims with separate evidence, so the citation on each record page states them as two separate blocks, followed by a note saying exactly how far that record was processed.

The snapshot itself is dated, not versioned. No semantic version is invented for a corpus with no release discipline: use the snapshot date. When the corpus changes, the date advances and the change is recorded in the changelog below.

Citing the corpus

Indo-Pacific Record. China Desk Corpus. Snapshot — 2026-08-26. 3,574 records. Benjamin Yang, Creator and Editor.

Citing a record

Each record page carries its own citation in two blocks — the source text as its publisher issued it, then the record as this corpus holds it — followed by a note stating exactly how far that record was processed. The source-text block always quotes the original-language title; the machine translation is never substituted for it.

Citing an edition

Editions of The PLA Watch are published artifacts with their own issue numbers and canonical addresses, and they are cited as published. On Analysis, each edition in the archive carries its own citation behind a “Cite this issue” disclosure. Week-ending dates read as 8 August 2026 in citations, while the archive table keeps ISO dates.

Corpus changelog

A hand-written record of what changed in this corpus and what is known to be missing from it. It is maintained by the editor, entry by entry. It is not generated from repository history, and it is not a collection-health log — current run results and the vocabulary used to describe them belong to Coverage, which this does not duplicate.

Initial documented snapshot — 2026-08-26

The first snapshot of this corpus to be documented for readers. What follows begins with what is missing from it.

A recorded collection interruption
No pipeline run is recorded on the UTC dates 2026-07-17 through 2026-07-24, and no record in this snapshot carries a source-stated publication date inside that window.
The interruption cannot be quantified
What those dates would have held is not recoverable from stored data. Run records preserve no target publication date and no source-level outcome, so the volume missed can be neither reconstructed nor estimated. It is not counted, and its absence is not evidence that nothing was published.
A processing backlog
837 records have not completed analysis and are not settled as out of scope: 834 are awaiting screening and 3 passed screening without a completed analysis.
Most records carry no English rendering
2,239 of 3,574 records hold no machine-translated title. The same records — the same identifiers, not merely the same total — also hold no machine summary and no analysis-model record, and they are exactly the records outside the Analyzed state. English coverage is a property of the analyzed subset, not of the corpus.
Some captures stored no source text
45 records hold an empty body capture. The original URL remains recorded for each; the empty value is a stored capture defect, not evidence that the page was blank.
Repeats across outlets are held once
De-duplication across sources is first-writer-wins. An item carried by more than one outlet is stored under whichever source reached it first, so per-outlet totals count stored records rather than how often something was published.
Snapshot and size
This snapshot is dated 2026-08-26 and holds 3,574 records from 4 outlets and 3 institutions, with source-stated publication dates from 2026-05-07 onward.