Data landscape report

The State of Chinese Case Law Data, 2026

If you are building a legal AI product with any China exposure—cross-border disputes, IP enforcement, investigations—your retrieval layer eventually hits the same question: where does mainland Chinese case law actually come from, and how much of it can you still get? In 2026 the answer is less obvious than it was five years ago, and considerably more consequential.

China's courts have published the largest body of judicial decisions in the world: roughly 143 million effective judgments accumulated on China Judgements Online as of December 2023, according to figures cited in Supreme People's Court (SPC) press briefings. But annual publication fell from 19.2 million documents in 2020 to about 5.11 million in 2023 before partially recovering, and since 2024 the system has split into an internal full-text database and a small public library of curated cases.

This report assembles the verifiable numbers—publication volumes year by year, the 2024 restructuring of judicial transparency infrastructure, what Western platforms list for China coverage, and the funding race among legal AI vendors—into one picture for teams deciding how to source Chinese case law data in 2026.

The world's largest case law corpus, by the numbers

China Judgements Online (CJO) is the SPC-operated public repository of effective court decisions. By December 2023 it had accumulated approximately 143 million published judgments, and cumulative site visits exceeded 108.1 billion—figures cited by SPC officials in a December 2023 press briefing. No other jurisdiction publishes court decisions at anything close to this scale.

The catch: cumulative size says little about forward availability—which is where the curve comes in.

Chinese court data availability, 2020–2025: the full curve

Annual publication on CJO tells a clear story: a 2020 peak, a steep three-year decline, then a partial rebound. The figures below are drawn from SPC press briefings (the 2023 number was given by SPC officials in December 2023) and from the SPC work report delivered in March 2026.

YearJudgments published on CJOShare of 2020 peak
202019.2 million100%
202114.9 million~78%
202210.4 million~54%
2023~5.11 million~27%
20249.69 million (+92.7% YoY)~50%
2025~10.98 million~57%

Read it both ways. The optimistic reading: 2024 and 2025 both grew, and the March 2026 work report put 2025 publication at 10.979 million documents, the strongest year since 2021. The sober reading: after two years of recovery, inflow still runs at about 57% of the 2020 peak. Simple arithmetic shows the depth of the trough—at 2020 rates, 2021–2023 would have added roughly 57 million documents; the actual total was about 30 million.

The 2024 restructuring: one internal database, one public library

In January 2024 the National Court Judgments Database went live: a comprehensive full-text repository searchable only on the court system's internal network. Judges can query it; lawyers, academic researchers, and the public cannot.

A month later, on February 27, 2024, the SPC opened the People's Court Case Library to the public. It is a different kind of resource: a curated set of reference cases reviewed and approved by the SPC to guide adjudication. It launched with 3,711 cases and had grown to roughly 5,327 by January 2026.

For anyone tracking China judicial transparency data, the two tracks serve different purposes—and neither restores what CJO offered at peak:

What the availability curve means for legal AI builders

1. Historical snapshots are the scarce asset. Annual inflow can recover; history cannot. A corpus assembled across the full publication window—everything posted before and through the trough—cannot be reconstructed from today's feed. Teams evaluating Chinese case law data in 2026 are really evaluating the depth and integrity of a snapshot, not the freshness of a stream.

2. The 2021–2023 vintage is disproportionately rare. Decisions from those years entered the public record at roughly 78%, 54%, and 27% of peak rates respectively. Any corpus crawled recently, or built mainly from the post-2024 rebound, will be thinnest in exactly that window. For retrieval products this surfaces as missing precedent; for fine-tuning and evaluation, as a skewed temporal distribution.

3. Citation-grade products need primary documents, not digests. RAG pipelines that must show a verifiable source, extraction models that need full procedural histories, and benchmarks that test reasoning over real holdings all require complete judgment text with stable identifiers: case number, court, date, cause of action, outcome. Summaries and translated selections support reading; they do not support grounding.

The global legal AI race is now a content-licensing race

Capital flows explain why jurisdiction-by-jurisdiction data sourcing is now a board-level topic in legal AI. In March 2026, Harvey raised $200 million at an $11 billion valuation; in April 2026, Legora reached a $5.6 billion valuation with annual recurring revenue past $100 million. Both announced new Singapore and Tokyo offices in 2026—the Asia-Pacific push is explicit. In November 2025, Clio acquired vLex for $1 billion—a corpus spanning, by vLex's own count, 110 jurisdictions and more than 1 billion documents.

Public announcements show two patterns in how these platforms acquire law:

Against that backdrop, mainland China stands out as the largest body of case law still largely absent from Western research stacks. As of mid-2026, vLex's public China coverage page lists international collections and secondary commentary, and mainland China does not appear in its primary-law jurisdiction list; LexisNexis's dedicated Chinese law file (LNCHNL) is described as roughly 3,000 selected documents in English translation; the Westlaw China site was not accessible when checked. That is not a criticism; it is a measure of how hard the jurisdiction is to cover—and of where the per-jurisdiction licensing playbook has yet to be applied.

Chinese text is scarce where models are trained

There is a second scarcity argument, independent of legal coverage. English accounts for roughly 59.8% of content on the web; Chinese accounts for about 1.3%—a share that has fallen from 4.3% over the past eleven years. High-quality, formal, domain-specific Chinese text is dramatically underrepresented in the open data that general-purpose models train on.

Court judgments are among the best Chinese text available for closing that gap: standardized structure, formal register, dense statutory citation, consistent metadata. A judgment corpus is therefore valuable twice over—as legal ground truth for China-facing products, and as scarce, high-quality Chinese language data for any multilingual legal model.

Compliance basics for cross-border licensing (informational, not legal advice)

The following is informational background, not legal advice. Any cross-border data licensing deal should be structured with qualified counsel.

Chinese court judgments are published under SPC rules that retain party names while redacting identifiers such as ID numbers and home addresses. Around that baseline, three features of the Chinese data regime shape how licensing deals are typically structured:

The practical takeaway: ask any vendor to explain its sourcing basis, its PIPL position, and exactly which uses the license excludes—then verify with counsel. We cover the framework in more depth in our compliance primer on PIPL, the Data Security Law, and cross-border transfer of Chinese court data.

Three recommendations for data buyers in 2026

If your team is evaluating Chinese case law data this year, the numbers in this report suggest three concrete tests.

  1. Buy the snapshot, verify the vintage. Ask every vendor for a coverage breakdown by year, court level, region, and case type—and look hardest at 2021–2023 density, the window in which public inflow fell to about a quarter of peak. A corpus that is thin there stays thin; incremental updates do not fill a historical gap.
  2. Test machine-readability on your own pipeline before you sign. Structured fields—case number, court, date, parties, cause of action, outcome—matter more than headline document counts. Run a trial API against your real retrieval and extraction workloads, measure parse and dedup behavior, then negotiate.
  3. Put scope and updates in the contract. Pin down the sourcing basis, permitted uses and exclusions (including foreign-discovery carve-outs), and terms for incremental updates beyond the base snapshot. Our guide to licensing Chinese court judgment data walks through deal structures, delivery formats, and evaluation checklists in detail.

SinoVerdict's offer maps directly onto these tests: a structured, machine-readable snapshot of 170M+ Chinese court judgments with coverage through 2023, delivered as a bulk dataset, REST API, and MCP server, with incremental updates available on request.

This article is informational only and does not constitute legal advice. Figures are drawn from publicly reported sources, including Supreme People's Court press briefings and work reports.

Frequently asked questions

How many Chinese court judgments are publicly available?

China Judgements Online, the Supreme People's Court's public repository, had accumulated approximately 143 million effective judgments as of December 2023, according to figures cited in SPC press briefings. It is the largest public collection of court decisions in the world, and cumulative visits to the site exceed 108.1 billion.

Is China Judgements Online still being updated in 2026?

Yes, but well below peak rates. Annual publication fell from 19.2 million documents in 2020 to about 5.11 million in 2023, then recovered to more than 9.69 million in 2024 and roughly 10.98 million in 2025, according to figures cited in SPC press briefings and the March 2026 SPC work report. The 2025 volume is still only about 57% of the 2020 peak.

What is the difference between the National Court Judgments Database and the People's Court Case Library?

The National Court Judgments Database, launched in January 2024, is a comprehensive full-text repository available only on the Chinese court system's internal network, closed to lawyers, researchers, and the public. The People's Court Case Library, opened to the public on February 27, 2024, is a curated set of reference cases approved by the Supreme People's Court; it launched with 3,711 cases and held roughly 5,327 as of January 2026.

Do Western legal research platforms cover mainland Chinese case law?

Coverage is limited as of mid-2026. vLex's public China coverage page lists international collections and secondary commentary, and mainland China does not appear in its primary-law jurisdiction list; LexisNexis's dedicated Chinese law file is described as roughly 3,000 selected documents in English translation; and the Westlaw China site was not accessible when checked. Teams that need full-corpus primary law generally source it through specialized data licensing.

Can Chinese court judgment data be licensed for AI training and retrieval?

Yes. Judgments are published under Supreme People's Court rules, and cross-border licensing is typically structured around China's Personal Information Protection Law and Data Security Law, which provide defined pathways for data export. Licenses commonly exclude certain uses, such as providing data to foreign judicial or law-enforcement authorities, which Article 36 of the Data Security Law restricts. This is informational background rather than legal advice, so structure any deal with qualified counsel.

What does SinoVerdict provide?

SinoVerdict licenses a structured, machine-readable corpus of more than 130 million Chinese court judgments, with coverage through 2023 and incremental updates available on request. Delivery options are a bulk dataset, a REST API, and an MCP server. A trial API key and a corpus coverage report are available on request for evaluation.

Verify the coverage before you build on it

Request a coverage report to see how SinoVerdict's 170M+ judgment corpus breaks down by year, court level, region, and case type — or apply for a trial API key and test the corpus against your own retrieval and extraction workloads.

Request trial access