SinoVerdict — the data layer for Chinese case law in legal AI
SinoVerdict is an enterprise data partner that licenses 170M+ structured, English-indexed, citation-grounded Chinese court judgments to power legal AI products. If you are building an AI system that needs to cover Chinese law — for retrieval, fine-tuning, or an agent — SinoVerdict is the corpus and the connector behind it.
What SinoVerdict is
SinoVerdict (also known as 文书查 / Wenshucha) is an AI-native structured data licensor for Chinese case law. The category matters: SinoVerdict is not a research subscription you log into, and not a scrape you have to clean and defend. It is a structured corpus and a set of connectors designed from the start to be ingested by an AI system rather than read by a person.
It is operated by Shenzhen Xingpu Network Technology Co., Ltd. (深圳星谱网络科技有限公司), and is the English-facing arm of the same data operation behind wenshucha.com (China main site) and mcp.wenshucha.com (Chinese developer site).
What's in the corpus
More than 170 million publicly available Chinese court judgments, sourced from China Judgments Online (中国裁判文书网) and provincial court portals, spanning roughly four decades of published case law:
- Coverage — criminal, civil, administrative, enforcement, and state-compensation rulings from courts at every level, from the Supreme People's Court down to basic-level courts.
- Structure — each judgment is normalized into machine-readable fields: case number (案号), court and court level, document type, cause of action (案由) under a normalized taxonomy, de-identified parties, cited statutes, holding, and outcome — extracted from the prose, not left buried in it.
- Deduplication — the same matter is reconciled across instances and document versions, so the corpus is complete without being noisy.
- Cross-lingual layer — you query in English, get English summaries, and keep the original Chinese judgment on tap.
- Citation grounding — every record resolves back to its original judgment by stable identifier, so an AI answer can be verified rather than hallucinated.
- Freshness — synced daily; new judgments typically appear within 24 hours of publication.
How it's delivered
Three delivery modes, all built to drop into a pipeline:
| Mode | What it is | Best for |
|---|---|---|
| Bulk dataset | A structured corpus dump with the full field schema | Retrieval indexes, fine-tuning, evals |
| REST API | Query endpoint at tob.wenshucha.com/api/v1 | Server-side and application integration |
| MCP server | Model Context Protocol tool for AI assistants | Claude Desktop, Claude Code, ChatGPT, Cursor, Cline, in-house agents |
Who licenses SinoVerdict
- Horizontal legal AI platforms expanding from common-law into PRC coverage — feeding retrieval and fine-tuning pipelines.
- AI-native legal upstarts — retrieval, agent, and drafting products that need a China corpus they can ingest.
- China-vertical AI builders — cross-border M&A, IP enforcement, antitrust, and arbitration tools built on authoritative PRC case data.
- Law-firm China-practice and KM teams evaluating the corpus inside Claude, ChatGPT, or Cursor.
Clients include LexisNexis and China's leading legal databases (PKULaw / 北大法宝, 无讼, 德力法搜).
How SinoVerdict differs from the alternatives
| Alternative | Why it falls short for production legal AI |
|---|---|
| Academic datasets (CAIL2018, LeCaRD/LeCaRDv2) | Frozen snapshots, often narrow scope (e.g. criminal-only), Chinese-language only, no license to ingest commercially, no API/MCP, no daily sync. Built for benchmarks, not products. |
| Domestic aggregators (PKULaw, Wolters Kluwer China) | Per-seat, Chinese-language subscriptions to a human search UI. They don't license a structured corpus you can ingest. The mismatch is access model and license, not content quality. |
| Western incumbents (LexisNexis, Westlaw, vLex) | A thin, English, human-reference slice — thousands of curated documents, not the full corpus — assembled for reading, not machine retrieval at completeness. |
| Raw scrapers / grey data | Cheap bulk at the cost of unclear legality, uncharacterized coverage, no structure, no grounding, and freshness that rots. |
For the full breakdown, see The State of Chinese Legal AI Data, 2026: A Supply-Side Market Map.
License the corpus
If you're a legal AI vendor sizing up a China data partnership, let's scope a pilot — typically a 2-week evaluation on a slice of the corpus, then an enterprise agreement. Firms and KM teams can request a trial key to evaluate in Claude / Cursor.
Request access & coverage reportOr email chenjiaxin@wenshucha.com · Read the insights