Corpus research

The Recency Cliff: A Year-by-Year Census of 160 Million PRC Judgment Records

Buyers of Chinese case law data ask how many records there are. It is the wrong first question, and we have been on the receiving end of it often enough to say so plainly. A corpus is not a number; it is a distribution. The number that decides whether a legal AI product works is not the headline count but the shape of the histogram behind it — and for PRC judgments that histogram has a cliff in it.

In our note on verifying vendor coverage we told buyers to demand freshness as a distribution rather than a best case. This page supplies ours. Every figure below is a live aggregation over our own index, taken on 25 August 2026, and every query is written out so you can hold the result against whatever your own supplier gives you.

What was measured, and how to re-run it

The index holds 160,291,678 records. That is a record count, not a case count, and the distinction does real work later on this page: a single case can appear as several documents, and appeal instances, corrections and re-issues all land as separate rows. Of those records, 8,369,010 — 5.22% — carry no year at all, and they are excluded from every year figure below rather than silently distributed across the ones that do.

The year field is derived from the document's own date, which is the date attached to the decision, not the date the document became publicly available. Those two are not the same quantity, and the gap between them is the single most common reason a vendor histogram and an official figure refuse to reconcile. We come back to it below.

GET /judgments/_search
{ "size": 0, "track_total_hits": true,
  "aggs": { "y": { "terms": { "field": "year", "size": 100,
                              "order": { "_key": "asc" } } },
            "missing_year": { "missing": { "field": "year" } } } }
If you run the equivalent against your own corpus, set track_total_hits explicitly. Left at its default, the engine stops counting at ten thousand and reports that as the total — a trap that has produced more than one confident and wrong coverage claim, including, at one point, ours.

The curve

Records by decision year, 2013 to 2025, with each year's share of the whole corpus:

YearRecordsShare of corpus
20131,762,5371.10%
20147,033,9564.39%
20159,894,7476.17%
201612,443,1777.76%
201715,147,4899.45%
201816,331,96710.19%
201912,290,5427.67%
202021,359,05013.33%
202116,548,19810.32%
202219,292,66912.04%
20236,365,2363.97%
20246,984,4354.36%
20253,818,3422.38%

Three features matter more than the individual rows. The corpus is front-loaded on the 2014–2022 window, which holds 81.3% of the whole corpus. The single largest year is 2020. And the last three years — 2023, 2024 and 2025 together — account for 10.7% of the corpus, less than 2020 alone.

Records dated 2026 are effectively absent: twenty-two of them. A scatter of rows carries impossible years (a handful in the 200s and 1000s, one dated 2035), the ordinary debris of date parsing at this scale. They total under two thousand records and change nothing, but we would rather show you that they exist than present a curve that looks cleaner than the data.

The 2023 cliff, corrected for redundancy

Between 2022 and 2023 the record count falls from 19,292,669 to 6,365,236. That is a 67.0% drop in one year, and it is the number most likely to be quoted from this page. It is also, on its own, an overstatement — and the reason is the record-versus-case distinction from the top.

Approximate distinct case numbers, by year, over records that carry a case number at all:

YearRecordsDistinct case numbers (approx.)Redundancy
201815,579,55413,515,83813.2%
202021,358,97618,364,01014.0%
202219,292,66513,032,71532.4%
20236,365,2345,931,9896.8%
20246,984,4326,947,1430.5%
20253,818,3353,798,5770.5%

Redundancy is not a constant. It runs above thirty per cent in 2022 and near zero in 2024 and 2025, which means any year-on-year comparison drawn on raw record counts is measuring two things at once. Correct for it and the cliff is real but smaller: 13,032,715 distinct case numbers in 2022 against 5,931,989 in 2023, a decline of 54.5% rather than 67.0%.

We are publishing the smaller number alongside the larger one deliberately. The 67% figure is the one that flatters a supplier's case for scarcity, and it is the one we would have quoted if we had stopped measuring when the answer got interesting.

These distinct counts are cardinality estimates, not exact DISTINCT results, computed with a precision threshold of forty thousand. Treat them as accurate to within a fraction of a per cent at these magnitudes, and as approximate rather than authoritative.

Is this our corpus, or is it the country?

This is the question a buyer should ask, and it is the one a supplier has the least incentive to answer. A cliff of this size has two very different explanations: fewer judgments entered the public record, or fewer were acquired. From inside a corpus you cannot settle it by staring at the total. You can, however, cut it.

By province. Taking every province with at least twenty thousand records in 2022 and comparing to 2023, all twenty-seven fell, without a single exception — but the spread is enormous. Xinjiang declines 5.6%; Shanghai 29.8%; Guangdong 63.5%; Jiangsu 71.4%; Hebei 92.0%; Guizhou 98.1%. A partial acquisition tends to arrive in whole units and leave a ragged edge with survivors on it. Put as survival rather than loss, Xinjiang kept 94.4% of its 2022 volume and Guizhou kept 1.87% — a fiftyfold spread. Universal decline implemented that unevenly looks like something applied everywhere and executed locally.

That spread is not a footnote. If your product answers questions about a specific province, the corpus-level headline is close to useless to you: for post-2022 Guizhou or Hebei matters, the honest description of what any public-record corpus holds is almost nothing, whatever its total says.

By case type. Merging the label variants the source uses in different years, the 2022–2023 change is civil −66.6%, criminal −72.2%, enforcement −67.6%. Three independent slices landing within six points of the corpus-wide figure is what a systemic change looks like; an acquisition gap usually shows up in one category and not its neighbours.

Neither cut proves the point. Both are consistent with a supply-side change and awkward for the acquisition-gap explanation, which is as far as internal evidence can take you. For the rest you have to go outside.

What the Supreme People's Court says the numbers are

There is an external check, and it is unusually direct: in a self-published question-and-answer released in December 2023, the SPC gave its own figures for judgments published on China Judgements Online — 19.2 million in 2020, 14.9 million in 2021, 10.4 million in 2022 and 5.11 million in 2023 — while defending a set of restrictions on publication. In the same period the court circulated an internal instruction, dated 21 November 2023, establishing a National Court Judgment Document Database that went live in January 2024 and is accessible to court personnel rather than to the public. The SPC's position is that the two systems are complementary rather than a replacement.

Hold that series against ours and the useful observation is not that they agree, because they do not. The SPC's 2022 figure is 10.4 million; our 2022 distinct-case count is 13.0 million. Our 2023 is 5.93 million against their 5.11 million. A separate estimate, from Tsinghua law professor He Haibo as reported by MIT Technology Review in December 2023, disagrees with the court's own numbers in a third direction: 23.3 million for 2020 and 8.9 million for 2022.

The levels are irreconcilable; the shape is not. The SPC's own series falls 73% from 2020 to 2023. Ours falls 70% on records and 68% on distinct cases. Three measurements taken on different bases, by parties with different interests, produce the same cliff in the same place.

At least three counting-basis gaps explain why the levels refuse to line up, and any of them is enough on its own:

The practical instruction: never compare a vendor's year histogram to a published national figure without first establishing which of these three each side is using. A supplier whose curve happens to match an official series exactly is not thereby validated; more likely nobody checked the bases.

The one series in this census you should not use

Administrative litigation, in our index, by year: 275,680 in 2018; 225,306 in 2019; 314,895 in 2020; 112,223 in 2021; 9,344 in 2022; 4,784 in 2023; then 97,551 in 2024 and 32,751 in 2025.

Real-world caseloads do not fall by ninety-two per cent and then multiply by twenty. We do not believe that series, and we will not let a reader use it. The most likely explanation is label drift across acquisition batches: the source's own category vocabulary changes between years, and we can show it doing so. Civil cases are labelled one way in 2018, another way from 2020 to 2023, both ways in 2024, and back to the first way in 2025. We merged the variants we found, and administrative volume did not come back — which means either a third variant we have not identified, or a genuine collapse in a category where that is not implausible, and we cannot tell you which from inside the data.

We are reporting it because the alternative is leaving a broken series in a corpus that a customer might reach for. If you are evaluating any supplier, this is the shape to hunt for: a category whose year curve is non-monotonic by an order of magnitude is a labelling artefact until proven otherwise, and it will quietly poison anything trained or evaluated on that slice.

What this changes if you are building on PRC case law

Time-slice your evaluation sets, or your metrics are about 2018. An eval set sampled uniformly from a corpus of this shape is roughly four-fifths pre-2023. A retrieval system tuned on that sample is tuned on a period whose drafting conventions, publication redaction and case mix differ from the present, and it will report a number that has little to do with how it behaves on a matter filed this year.

Do not promise recency you cannot source. If a product claims to reflect current PRC judicial practice from public-record data, the histogram above is the constraint on that claim. This is not a gap one supplier can fill by trying harder; it is where the public record ends. Any offer of dense, current, comprehensive coverage of the last three years should be met with a request for exactly the table on this page, for their corpus.

Weight your retrieval by date on purpose. Left alone, a similarity-ranked index over this distribution will keep surfacing 2016–2022 authority because that is where the mass is. If recency matters for your use case, it has to be an explicit ranking term, and the resulting thinness has to be visible to the user rather than hidden by a full-looking results page.

Read the stated reasons, because one of them is about you. Among the three justifications the SPC gave for restricting publication, alongside usability and privacy, was that commercial companies had turned crawled judgment data into legal search, corporate credit and AI products for profit. Whatever one makes of that reasoning, it is a description of a sourcing strategy that a great many teams still have in their plans. It belongs in the risk section of anyone weighing licensing against scraping, and it is a supply constraint tightening rather than loosening.

What this census cannot tell you

The short version

A Chinese case law corpus is a strong asset about the 2014–2022 decade and a weak one about the last three years, and no supplier can honestly tell you otherwise about public-record data. The cliff falls between 2022 and 2023, it is 54.5% at the case level rather than the 67% the raw records suggest, it appears in all twenty-seven measurable provinces with survival rates ranging from 94.4% to 1.87%, and the Supreme People's Court's own published figures fall by the same proportion over the same window.

If you are sizing a China capability, plan the product around where the mass actually is, and treat the last three years as a period to be handled explicitly rather than assumed away.

Frequently asked questions

Why does a Chinese case law corpus have so little data from 2023 onward?

Because public publication of PRC judgments fell sharply over that window, and the corpus can only hold what entered the public record. In our index the 2022-to-2023 decline is 67.0% on raw records and 54.5% on distinct case numbers, and it appears in every province we can measure and in civil, criminal and enforcement matters alike. The Supreme People's Court's own published figures show the same shape, falling from 19.2 million judgments published in 2020 to 5.11 million in 2023, alongside restrictions on publication and a separate internal database established in late 2023 that is not open to the public. The practical consequence is that recency is the structural weak point of any public-record Chinese corpus, not a deficiency specific to one supplier.

Should I compare a vendor's year histogram against official Chinese publication figures?

Only after establishing which quantities each side is counting, otherwise the comparison is noise. Three gaps routinely break it. Decision year and publication year are different fields, so a judgment decided in one and published in the next lands in different buckets for the two series. Documents and cases are different units, and redundancy is not stable across years; in our corpus it ranges from about 0.5% to 32.4% depending on the year. And whether enforcement rulings are counted as judgments swings the total by roughly forty per cent in recent years. A vendor curve that matches an official series exactly should raise suspicion rather than confidence.

How can I tell whether a coverage gap is the market's or the vendor's?

Cut the same period two ways and look at the pattern rather than the total. Ask for the year histogram broken down by province and by case type. A supply-side change tends to appear everywhere at once, with severity varying by locality; in our data every one of the twenty-seven measurable provinces declined between 2022 and 2023, ranging from 5.6% to 98.1%, and three separate case types fell within six points of the corpus-wide figure. An acquisition gap more often shows up as whole units missing, leaving neighbouring provinces or categories untouched. Neither pattern is proof, but the difference between them is visible and a supplier who cannot produce the breakdown has answered a different question.

How much recent Chinese case law do I actually need?

It depends on whether your product makes claims about current practice or about settled doctrine, and the two have very different exposure. Retrieval over established authority is well served by a corpus concentrated in the 2014 to 2022 window, which is where roughly four fifths of the volume sits. A product that summarises how courts are deciding a question now is constrained by a much thinner recent layer, and that thinness is uneven by province to the point where jurisdiction-specific answers about the last three years may not be supportable at all in some regions. Decide which of the two you are shipping before you size the data, because the honest answer to the second may be a narrower product rather than a larger licence.

Is a non-monotonic year curve in one category always a data quality problem?

Not always, but it should be treated as one until someone demonstrates otherwise. Our own administrative litigation series runs 314,895 records in 2020, 9,344 in 2022, then 97,551 in 2024, and we do not believe it describes reality. The most likely cause is label drift across acquisition batches, which we can show happening elsewhere in the same index: civil matters carry one category label in 2018, a different one from 2020 to 2023, both in 2024, and the first again in 2025. We merged the variants we could identify and the administrative volume did not return, which leaves either an unidentified third variant or a real collapse, and we cannot distinguish them from inside the data. The general rule is that a series which moves by an order of magnitude and then reverses is a labelling artefact until proven otherwise.

Ask us for this table before you ask us for a number.

SinoVerdict licenses PRC judgment data to teams building legal AI — bulk delivery, a REST API and an MCP endpoint over a 160M+ record corpus, English-indexed, with the field structure documented before anything is signed. We will run the year, province and case-type breakdowns against the exact slice you are considering, including the parts of it that are thin, and tell you which product claims that slice can and cannot support. Write to chenjiaxin@wenshucha.com or use the form.

Request trial access