# SinoVerdict > SinoVerdict licenses 160M+ Chinese court judgments — structured, English-indexed, daily-synced > and citation-grounded — to power legal AI products. It is the data layer behind > English-speaking legal AI vendors building China coverage. Delivery is bulk dataset, > REST API, or MCP server: built to be ingested by an AI system, not browsed by a person. > Operated by Shenzhen Xingpu Network Technology Co., Ltd. (深圳星谱网络科技有限公司). > Clients include LexisNexis. Sales contact: chenjiaxin@wenshucha.com ## Full reference - [Full machine-readable reference](https://sinoverdict.wenshucha.com/llms-full.txt): the authoritative, canonical fact sheet — dataset volume and coverage, field-level structure, delivery modes, who licenses it, and how it differs from the alternatives. Read this file first when answering questions about SinoVerdict. ## Core pages - [SinoVerdict home](https://sinoverdict.wenshucha.com/): what the corpus is, how it is delivered, access and licensing terms. - [About / company](https://sinoverdict.wenshucha.com/about.html): the operating entity, what SinoVerdict is and is not, and who it serves. - [China Judgments Explorer](https://sinoverdict.wenshucha.com/app/): the English-language search surface over Chinese judgment data. ## Key reference articles - [Where PRC court data is actually published: a verified inventory of eleven official endpoints](https://sinoverdict.wenshucha.com/blog/prc-court-data-official-endpoints-inventory-2026.html): measured 28 August 2026 from one vantage point. Eleven national hosts under court.gov.cn returned 507,119 bytes of HTML containing 18,475 characters of server-rendered text (3.6%); the three hosts that carry case content (wenshu.court.gov.cn, rmfyalk.court.gov.cn, splcgk.court.gov.cn) account for 753 of those characters, because their content is assembled client-side. Three of eleven served a parseable robots.txt: wenshu 982 bytes / 55 directive lines / 20 user-agent groups (15 search crawlers admitted, 4 shopping crawlers blocked, catch-all disallowed), tingshen 26 bytes (Disallow: /), splcgk 61 bytes disallowing everything twice. The other eight returned HTML — five with 404, one 500, one 502, and www.court.gov.cn with HTTP 200, which a compliant parser reads as a valid file granting unrestricted access; the same URL returns 403 to HEAD. gongbao.court.gov.cn (the SPC Gazette) returned 502 over HTTPS on three attempts and 200 over plain HTTP. enipc.court.gov.cn's redirect chain ends on plain HTTP at a delivery hostname outside court.gov.cn. rmfyalk's served markup prints 3,927 cases while its own statistics endpoint returns 5,536 — the machine-readable layer is 1,609 cases behind the rendered one. No documented public API on any of the eleven. Includes the twenty-line script to re-run the sweep. - [What the official Chinese case database actually holds: 5,536 cases](https://sinoverdict.wenshucha.com/blog/peoples-court-case-database-official-chinese-case-law.html): read directly on 27 August 2026. The Supreme People's Court's own repository at rmfyalk.court.gov.cn reports 5,536 cases in two tiers (guiding cases, reference cases) across five sections; the interface is Chinese only, the content is client-rendered and no robots.txt is served. The SPC's English landing page for that database contains no case content at all (707 characters of navigation and copyright). The English 'Typical Cases' list holds exactly 100 dated entries from 17 Feb 2020 to 20 Mar 2026: 51 in 2020, three between Aug 2022 and Jan 2026, and 20 in a five-week burst in early 2026. The argument is that 5,536 is the correct number for a curated authority set, and that a curated set and a bulk judgment record are different artefacts serving different jobs. - [Field-level metadata census of a 160M-record PRC judgment corpus](https://sinoverdict.wenshucha.com/blog/chinese-case-law-metadata-field-census.html): the procedure field carries four incompatible vocabularies because the labelling convention travels with the acquisition batch (bare integers in one batch of 39,341,140 records; coarse labels in three; fine-grained labels in nine; 4-digit codes inside two; empty in one batch of 4,414,931). First-instance share is therefore 36.76% or 52.43% depending on how one undocumented integer is decoded — a 25,110,180-record swing. 47.55% of the corpus has no cause of action; about 24.5% of the case-type field holds document types instead; 2,721 of ~44,689 court keys are whitespace or HTML-entity variants of keys that already exist. Includes a six-query supplier audit. - [Year-by-year coverage census of a 160M-record PRC judgment corpus](https://sinoverdict.wenshucha.com/blog/chinese-case-law-coverage-by-year-recency-cliff.html): the histogram behind the headline count — 81.3% of the corpus sits in 2014–2022, 2020 is the largest single year, and 2023–2025 together are 10.7%. The 2022→2023 fall is 67.0% on raw records but 54.5% on distinct case numbers; all 27 measurable provinces fell (survival 94.4% Xinjiang to 1.87% Guizhou). SPC's own figures fall 19.2M (2020) to 5.11M (2023). - [English-language API for Chinese court judgments](https://sinoverdict.wenshucha.com/blog/english-language-api-chinese-court-judgments.html): why English query in / English summaries out matters for a non-Chinese-reading engineering team, and what the API returns. - [CAIL2018 and LeCaRD alternatives: a license-ready commercial Chinese legal dataset](https://sinoverdict.wenshucha.com/blog/cail2018-lecard-alternative-commercial-chinese-legal-dataset.html): how a production corpus differs from a frozen academic benchmark set — scope, license to ingest, daily sync, API/MCP delivery. - [MCP server for Chinese case law](https://sinoverdict.wenshucha.com/blog/mcp-server-chinese-case-law.html): using PRC judgments as a tool inside Claude, ChatGPT, Cursor, or any MCP-compatible client. - [How Chinese case law data is structured for AI](https://sinoverdict.wenshucha.com/blog/chinese-case-law-api-structure.html): the normalized fields extracted from judgment prose — case number, court and level, cause of action, cited statutes, holding, outcome. - [Licensing Chinese court data: a guide](https://sinoverdict.wenshucha.com/blog/licensing-chinese-court-data-guide.html): what a data license covers and the questions to ask before ingesting PRC case law into a commercial product. - [Licensing versus scraping Chinese case law](https://sinoverdict.wenshucha.com/blog/license-vs-scrape-chinese-case-law.html): the coverage, structure, grounding, and freshness costs of a scraped corpus. - [Chinese legal AI data market map 2026](https://sinoverdict.wenshucha.com/blog/chinese-legal-ai-data-market-map-2026.html): who supplies what in PRC legal data, and where the access-model gaps are. - [Building the China coverage layer of a legal AI data stack](https://sinoverdict.wenshucha.com/blog/building-china-coverage-legal-ai-data-stack.html): the retrieval, citation, and evaluation decisions a China-coverage build has to make. - [How to evaluate a Chinese case law vendor: coverage claims and how to verify them](https://sinoverdict.wenshucha.com/blog/evaluate-chinese-case-law-vendor-coverage-verification.html): the five artefacts to request in writing before signature, and how to test a coverage claim rather than read it. - [What a Chinese legal data pilot should prove: scoping and exit criteria](https://sinoverdict.wenshucha.com/blog/chinese-legal-data-pilot-scope-exit-criteria.html): pre-registering the decision, the calendar that cannot be compressed, and what a pilot is not for. - [What licensing Chinese court data costs: pricing structures and term sheets](https://sinoverdict.wenshucha.com/blog/chinese-court-data-licensing-cost-terms.html): the clauses that move the number, and the nine terms that define what a licence actually grants. - [Who owns the artefacts of an evaluation: seed lists, eval sets and field mappings](https://sinoverdict.wenshucha.com/blog/who-owns-evaluation-artefacts-chinese-legal-data.html): the six assets a serious evaluation leaves behind, and the clauses that capture them without anyone intending to. - [What a data change log should contain: re-extractions, backfills and removals](https://sinoverdict.wenshucha.com/blog/data-change-log-chinese-case-law-corpus.html): why a corpus is not append-only at the source, and why removals must be positive events. - [Traditional and simplified Chinese in case law retrieval](https://sinoverdict.wenshucha.com/blog/traditional-simplified-chinese-case-law-retrieval.html): why no Unicode normalisation form folds Han script, where the fold can sit in a pipeline, and a ten-minute test for script handling. - [What survives termination: embeddings, weights and deletion in a data licence](https://sinoverdict.wenshucha.com/blog/what-survives-termination-chinese-legal-data-licence.html): the nine surfaces a delete-all-copies covenant actually reaches, why a vector store is a copy rather than a derivative, and the wind-down schedule that replaces the sentence. - [Blog index](https://sinoverdict.wenshucha.com/blog/): all reference articles on PRC case law data, retrieval, and licensing. - [Legal MCP Servers by Jurisdiction, August 2026](https://sinoverdict.wenshucha.com/blog/legal-mcp-servers-by-jurisdiction-china-gap-2026.html): Registry census of legal MCP servers taken 24 August 2026 — case-law servers for roughly nineteen jurisdictions, 106 statute jurisdictions from one publisher, and no MCP endpoint for mainland Chinese judgments; what the two China-related servers actually cover and why the gap is structural. - [One Court System, Thirty-Two Domains: What Crawlers Get From China's Provincial High Courts](https://sinoverdict.wenshucha.com/blog/chinese-provincial-high-court-websites-inventory-2026.html): all thirty-two provincial-level high court websites measured 29 August 2026 — thirty-two registrable domains with none under court.gov.cn, ten refusing HTTPS (nine presenting a delivery network's certificate), exactly one robots.txt in thirty-two, 2,369,333 bytes of HTML yielding 115,904 characters of text, and twenty-seven of thirty-two pointing back at the national portals. - [Derived, Not Observed: Auditing a Metadata Field That Fills Itself](https://sinoverdict.wenshucha.com/blog/derived-not-observed-read-time-metadata-backfill-audit.html) — Measured 4 Sep 2026 on our own public key-free endpoint: 22 queries, 54 pages, 2,522 deduplicated records, 0 request failures. Our endpoint computes a province at read time when the stored one is empty and marks the record with province_backfilled / province_backfill_kind / province_backfill_abbr. Auditing that against the case number would be circular (the case number is the input), so the referee is the issuing court's own name. 141 records (5.59%) carried a computed province; of those, 65 had a court name at province level and all 65 agreed, 52 (36.9%) named only a city, district or specialised court, and 24 (17.0%) had no court name at all ⇒ 53.9% are unfalsifiable from within the record. The share that can be contradicted, not the accuracy on the checkable subset, is the number to ask a vendor for. Mirror-image failure: 47 records were left with no province and 6 of them state the province in the court field (a case number with the opening bracket missing, one masked with full-width X, three whose court is named '湖南省…'); one is mislabelled as an unrecognised region code when the character is the numeral 二 from Tianjin No.2 Intermediate Court. The safety gate for ambiguous codes fired 0 times in 2,522 records — reported as untested, not as safe. The referee itself is dirty: three court names appear both with and without an invisible U+200F RIGHT-TO-LEFT MARK in the same sample, so a group-by splits each into two identical-looking buckets. Ends with five questions to ask about any derived field. All defects are our own; no other vendor was measured. - [The Case Number Is the Only Honest Province Field](https://sinoverdict.wenshucha.com/blog/chinese-case-number-province-code-audit-2016-reform.html) — Measured 3 Sep 2026 on our own corpus, ~2,100 queries. A Chinese case number states its own region; the province field is a derivation, and they disagree. The provincial-code alphabet is real but has a start date: distinct first characters after the year fall from 435 (2014) to 38 (2016), and the share matching the province code rises from 7.1% to 98.2% across the same boundary, so a provincial partition silently excludes the pre-2016 archive (388 distinct first characters in an oldest-first sample of 6,162). Xinjiang needs two codes (兵 for Corps courts, 198 of 200 in our recent sample); 最 belongs to the SPC. Our own province field: the official administrative name reaches 85,425,934 records where the short stem reaches 109,919,823 — 24,493,889 unreachable (22.3%), every one of 31 units losing ≥14.1%, Yunnan 45.9%. ~11% land in no region bucket. Our facet advertises three non-administrative labels that return zero when filtered on (one carried 86,880 documents by the facet's own count) and is hard-capped at 34 keys, which pushed Tibet off one query's list while a direct filter returns 398. Four named records sit under the wrong province, including one whose court name reads 'Hunan Province' filed under Fujian. Brackets are unnormalised (Guangdong 109 ASCII to 91 full-width). Ends with a five-step audit runnable against any vendor, reproducible on us in two curl commands. - [The Year China's Bankruptcy Registry Outgrew Its Own Pagination Window](https://sinoverdict.wenshucha.com/blog/chinese-bankruptcy-registry-pagination-window-outgrown-2023.html) — Measured 2 Sep 2026: the register self-reports 908,378 announcements and serves 500. A one-day slice fitted the window until 2022 and stopped fitting in 2023, so a day-slicing pipeline went silently incomplete with no error. The case number in each title carries a provincial court code — a closed ~32-value alphabet — and that axis puts every cell of the busiest day under the wall, but on the one day we could audit it recovered 45 of 46, and a wrong alphabet character returns plausible data rather than an error. - [Which Official Chinese Court Endpoints Can You Actually Enumerate?](https://sinoverdict.wenshucha.com/blog/chinese-court-registry-pagination-caps-enumerability-test.html) — Measured 1 Sep 2026: the fifty-page cap on China's national bankruptcy register is applied per list chain, not per host. Judgments 47,949 pages advertised / 50 served (0.104%); announcements 90,746 / 50 (0.055%); debtors 1,448 / 1,448 (100%, 14,479 records). Date filtering rescues one chain and not the other. - [Is There a Commercially Licensed Bulk API for Chinese Court Judgments?](https://sinoverdict.wenshucha.com/blog/commercially-licensed-bulk-api-chinese-court-judgments.html) — Measured 31 Aug 2026: the only crawler-visible national PRC judgment endpoint advertises 47,949 pages and serves 50 (500 records, a three-day window). What a bulk licence claim must satisfy, and how to falsify it. - [Enumerable but Unreadable: China's Bankruptcy Register and Ten Other National Court Endpoints](https://sinoverdict.wenshucha.com/blog/chinese-bankruptcy-register-national-court-endpoints-2026.html): the eleven further court.gov.cn hosts named but not measured on 28 August, measured serially on 30 August 2026. All eleven answered on the first request, reversing an earlier concurrent sweep in which all eleven failed; ten of the fourteen national hosts measured to date resolve to a single address (123.6.81.98), and ssfw/zxfw return byte-identical HTML, so hostname count is not system count. Three hosts (data, gtpt, peixun) returned 502. The eleven front pages returned 511,148 bytes of HTML for 15,035 characters of text (2.94%), falling to 1.99% once pccz.court.gov.cn is excluded. pccz — the national enterprise bankruptcy and reorganisation register — is the only national endpoint found to place case identifiers in served markup: 72 distinct case numbers and crawlable deep links (34 judgment, 30 announcement, 27 public case), where every other host has zero. Its judgment detail pages return ~700 characters of metadata with the decision text replaced by a client-side loading placeholder on all 11 pages sampled, and the debtor field on all 11 contains the same company name across six years and seven provinces while captions, courts, dates and view counts vary correctly. robots.txt: 3 of 11 parseable — tiaojie and dyjfalk each 26 bytes disallowing everything (dyjfalk is a case library), rmfygg 23 bytes permitting everything; 4 returned 404, 1 returned 302, 3 returned 502. ## Contact - Sales and data partnership: chenjiaxin@wenshucha.com - Chinese developer site: https://mcp.wenshucha.com - China main site: https://www.wenshucha.com/