Methodology and limitations
How a record gets here, and what it can and cannot support.
Scope
Indo-Pacific Record collects from declared desks, one jurisdiction or institutional family at a time. It does not monitor the Indo-Pacific comprehensively, and no page here should be read as claiming that it does. 1 of 4 declared desks collects today; the remainder are labeled on Desks with what specifically blocks each one.
- Records stored
- 3,574
- Records with stored original text
- 3,529
- Records with a machine translation of the title
- 1,335
- Records with a machine summary
- 1,335
- Collection runs recorded
- 122
- Publishing institutions
- 4
- Earliest item
- 2026-05-07
- Latest item
- 2026-08-26
Collection
Each source is declared in a desk manifest with its institution, language, listing endpoints and expected cadence. A scheduled run visits each listing, fetches items, extracts text, removes duplicates, and stores the original alongside metadata. Every source produces an explicit result each run, even when it produced nothing.
How current this is
Three dates, because they are three different facts and they come apart exactly when it matters. During a provider failure collection keeps running while analysis stops, and a single “last updated” line would have to pick one story to tell.
- Records last collected — 2026-08-26
- The newest document actually stored. Collection runs daily and is the part least likely to be interrupted.
- Analysis last produced — 2026-08-24
- The newest date on which any record was analyzed. It says when analysis last ran. It does not say that every record up to that date has been analyzed — records awaiting screening are counted separately on Coverage, and that count is the honest measure of the backlog.
- Last full update — 2026-08-24
- The last run that finished everything, including publication. A run that collects and then fails part-way never records one, so this date can sit behind the other two. When it does, the difference is the outage.
Any of the three may read not measured or not yet recorded. That is a real state — older runs did not record it — and it is printed rather than filled in with a plausible substitute.
Preservation and provenance
Each record stores the canonical URL it was retrieved from, the institution that published it, the source-stated publication date, the retrieval time, and a content hash computed once at insert. Original text, extracted text, machine translation, machine labels and human analysis are five separate layers and are labeled separately wherever they appear. A record's identifier never changes, and equal titles never collapse into one record — official institutions reuse titles for distinct events, and a title-only rule would erase the distinctions.
Translation and machine layers
Records are screened for relevance, translated, and summarized by language models. The model and prompt version are recorded per record. A machine translation is always shown as a translation beside the original, never in place of it. Machine classifications are labeled Model-flagged, never "significant".
What this archive supports
- That a named institution published a specific item on a specific date.
- That we retrieved it at a specific time, and what the original text was.
- Whether a source has published anything recently — and whether we successfully looked.
What it does not support
- That an institution believes what it published.
- That a described capability exists, or that an announced event occurred.
- That coverage is comprehensive. It is selective, heavily concentrated in one source, and regional only in declared scope — most of the region is not collected at all.
- That a machine-generated summary or translation is accurate without human review.
- Any quantified confidence in an analytical judgment. No calibration method exists, so no number is offered.
How evidence is labeled
These are the labels a reader will encounter:
- Source record the item as published, with its original language, canonical link and retrieval time.
- Official claim an assertion by the publishing institution — never rendered as established fact.
- Model-flagged a machine classification or summary, always marked as not reviewed by a human.
- Significant and Routine describe the editor's judgment of a weekly edition. They are separate from article-level Model-flagged classifications.
Known limitations
- Collection gaps are listed on Coverage and are not smoothed over.
- Some records are collected but not yet analyzed; they are counted, not hidden.
- One configured source has no working collector and is reported as such every run.
- Desks that do not collect show no record count. Zero would say the record was checked and found empty; the truth is that no record exists to check.
- Institutional authority is a claim about position, not truthfulness. A tier-A source is closer to the institution that speaks; it is not more likely to be accurate.
Field-by-field definitions, the identifier and citation rules, and the corpus changelog are in the Corpus Guide.