Chinese case law data, explained
Research, statistics and practical guides for teams building legal AI with Chinese court judgment data.
A surgical complication leaves a permanent impairment, a health regulator sanctions an institution or a licensed practitioner, or an implanted device is blamed for a revision operation — each is a China healthcare dispute, and each leaves a different kind of record. PRC medical damage liability runs on at least three tracks: fault-based civil claims against the medical institution under the Civil Code tort chapter (including the defined fault-presumption situations and the separate informed-consent duty to explain), an administrative track of health-regulator supervision and sanction against institutions and practitioners that produces decisions rather than judgments, and the drug and medical device overlap where product rules take over. Why the decisive artifact is the forensic appraisal referenced by the judgment rather than the adversarial expert testimony foreign teams go looking for; why liability arrives as a percentage of causal contribution rather than all-or-nothing, so risk models calibrated on binary outcomes systematically mis-price PRC exposure; and why the named defendant is the hospital — or the Chinese distributor or importer — rather than the doctor or the foreign brand, which breaks naive party-name searching. How a structured 130M+ judgment-and-ruling corpus — administrative decisions included — makes appraisal, causation, apportionment and sanction precedent findable in English, by track, claim type, specialty, court, year and disposition, with cited links to the originals.
A component fails and injures a user, a batch of imported supplements is found non-compliant with a national food safety standard, a regulator records a recall and penalises the domestic entity that placed the goods on the market, or a buyer sues the marketplace rather than the brand — each is China product liability, and each leaves a different kind of record. PRC product liability runs on at least three tracks: defect-based no-fault producer liability under the Civil Code, consumer-protection statutory multipliers (Consumer Rights Protection Law treble damages for fraud; Food Safety Law tenfold-price claims with a statutory minimum), and an administrative recall and penalty track run by market regulators that produces decisions rather than civil judgments — which is where recalls actually live, and which a judgments-only corpus systematically misses. Why foreign risk models misprice exposure by calibrating on compensatory logic when the decisive number is usually the multiplier and its base; why food and health products sit in a materially harsher regime than general goods; and why the foreign brand is frequently not the named defendant at all, because the importer, distributor or e-commerce platform is — which breaks naive party-name searching. How a structured 130M+ judgment-and-ruling corpus — administrative decisions included — makes defect, multiplier, platform-liability and recall precedent findable in English, by track, claim type, product category, defect type, court, year and disposition, with cited links to the originals.
When a discharge contaminates a watercourse and downstream users sue, when the ecology and environment bureau imposes a penalty and the operator challenges it, when a qualified environmental NGO or a procuratorate files a public interest action, when a government authority pursues ecological environment damage compensation, or when the same discharge is prosecuted as an environmental crime, the matter is China environmental litigation — and the defining feature is that the most consequential cases are often not brought by the injured party at all. This field is organised by standing, not subject matter. Why that precedent is plentiful yet almost unusable without structure: one pollution event can generate a penalty, an administrative challenge, private tort claims, an NGO or procuratorate public interest action, a government damage compensation claim and a prosecution; the private track runs on a reversed burden of proof on causation that foreign toxic-tort analogies invert; and the dispositive content is remedial and technical — restoration and substitute restoration orders, ecological damage appraisal figures, monitoring obligations — sitting in reasoning and, unusually often, in court-approved mediated settlements that a judgments-only corpus never sees. How a structured 130M+ judgment-and-ruling corpus — administrative decisions included — makes pollution tort, NGO and procuratorate public interest, ecological damage compensation, environmental administrative penalty and environmental crime precedent findable in English, by track, plaintiff type, remedy, court, year and disposition, with cited links to the originals — for cross-border environmental and ESG counsel, multinationals with PRC operations, and legal AI.
When the tax authority assesses additional corporate income tax after an audit, reclassifies a transaction for VAT or denies an input credit, makes a special tax adjustment on a multinational's related-party pricing, imposes a penalty for underpayment, or an underpayment tips into a criminal tax-evasion or fraudulent-invoicing case, the matter is a China tax dispute — and it is, for the most part, not decided on the ordinary commercial docket at all: the counterparty is the tax authority acting as an authority, most disputes over a tax amount must clear administrative reconsideration (often pay-first) before any court, and the case runs on the administrative-law track with a criminal edge. Why that precedent is plentiful yet almost unusable without structure: tax is not one dispute — it splinters into tax and claim types (corporate income tax, VAT and turnover taxes, transfer pricing and special tax adjustments, individual income tax, tax collection and penalties) that answer different questions, runs mostly on an administrative track behind a reconsideration-first, pay-first gate, keeps a decisive part of the record in reconsideration decisions and tax-authority documents rather than judgments, carries a criminal edge for evasion and fraudulent VAT invoicing on a separate track, turns on technical rules and circulars that change often, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus — administrative decisions included — makes corporate-income-tax, VAT, transfer-pricing, and collection-and-penalty precedent findable in English, by tax type, claim type, track, governing rule and date, authority or court, and disposition, with cited links to the originals — for cross-border counsel, multinationals, and legal AI.
When a departing engineer walks source code or a customer list to a competitor, a former employee breaches a non-compete, a rival copies a well-known product's name, packaging, or get-up, a company runs false advertising or defames a competitor, or a platform is accused of scraping a rival's data or hijacking its traffic, the matter is a trade secret and unfair competition dispute — and in China it does not run on one track at all: a rights holder can sue civilly for an injunction and damages, the market regulator can impose an administrative penalty for the same conduct, and serious trade-secret theft carries a criminal edge. Why that precedent is plentiful yet almost unusable without structure: unfair competition is not one dispute — it splinters into claim types (trade-secret misappropriation, employee non-compete and confidentiality, passing-off and confusion, false advertising and commercial defamation, commercial bribery, internet and data competition) that answer different questions, runs across civil, administrative, and criminal tracks, is decided on a technical finding (was the information secret and improperly accessed, is a get-up distinctive and confusing) buried in the reasoning, turns on regime-specific elements that evolve over time, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus — administrative decisions included — makes misappropriation, non-compete, passing-off, and internet-competition precedent findable in English, by claim type, track, governing rule and date, court, and disposition, with cited links to the originals — for cross-border counsel, technology companies, and legal AI.
When investors sue an issuer over a misleading prospectus or false disclosure, the securities regulator penalizes market manipulation or insider trading, a bond default drags in underwriters and other gatekeepers, or a controlling shareholder is pursued for a disclosure that moved a share price, the matter is a securities and capital markets dispute — and in China it does not run on one track at all: the decisive conduct finding is often made by the regulator in an administrative penalty, the civil investor-compensation action is built on that finding, and serious conduct carries a criminal edge. Why that precedent is plentiful yet almost unusable without structure: securities is not one dispute — it splinters into claim types (misrepresentation and investor compensation, market manipulation, insider trading, bond defaults and gatekeeper liability, fraudulent issuance) that answer different questions, runs across civil, administrative, and criminal tracks, anchors to a regulatory finding that lives outside the civil judgment, turns on regime-specific elements (materiality, loss causation, the damages measure, gatekeeper duties) that evolve over time, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus — administrative penalties included — makes misrepresentation, manipulation, insider-trading, and bond-default precedent findable in English, by claim type, track, governing rule and date, court, and disposition, with cited links to the originals — for cross-border counsel, investors, underwriters, and legal AI.
When Customs reclassifies goods under a higher-duty tariff heading, rejects a declared value, denies a claimed country of origin, imposes a penalty for under-declaration, denies an export tax rebate, or an under-declaration tips into a smuggling case, the matter is a customs and trade dispute — and in China it is, for the most part, not decided on the ordinary commercial docket at all: the counterparty is Customs acting as an authority, and the dispute runs on the administrative-law track, with a criminal edge for smuggling. Why that precedent is plentiful yet almost unusable without structure: customs is not one dispute — it splinters into claim types (tariff classification, customs valuation, rules of origin, customs penalties and smuggling, export rebates and processing trade) that answer different questions, is reviewed under administrative-law standards against an authority rather than as a contract, hinges on a technical determination (the heading, the valuation method, the origin finding) settled in administrative decisions and rulings rather than tidy civil judgments, clusters by Customs authority and reviewing court, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus — spanning the administrative track — makes classification, valuation, origin, penalty, and rebate precedent findable in English, by claim type, track, governing rule, court, and disposition, with cited links to the originals — for cross-border counsel, importers and exporters, customs brokers, and legal AI.
When cargo arrives short or damaged under a bill of lading, an owner and charterer fall out over hire and demurrage, two ships collide, a creditor arrests a vessel to secure a claim, a mortgagee enforces against a ship and competing maritime liens line up for priority, or a shipowner invokes limitation of liability, the matter is a maritime dispute — and in China it is heard not by the ordinary civil courts but by a dedicated system of specialized maritime courts, under a separate Maritime Code and a web of international conventions. Why that precedent is plentiful yet almost unusable without structure: maritime is not one dispute — it splinters into claim types (carriage and bills of lading, charterparties, collision and salvage and general average, ship arrest and maritime liens, ship finance and mortgages, marine insurance) that answer different questions, sits in its own forum and its own code, keeps its decisive steps (arrest, lien priority, limitation) in procedural rulings rather than judgments, turns on standard-form and convention wording, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus — rulings included — makes carriage, charterparty, collision, arrest, and lien-priority precedent findable in English, by claim type, governing rule, maritime court, and disposition, with cited links to the originals — for cross-border counsel, shipowners, cargo interests, P&I clubs, and legal AI.
When an insurer denies a claim for non-disclosure, an exclusion clause is contested, a liability policy is triggered by a product defect or construction accident, or an insurer that has paid pursues subrogation against the at-fault third party, the matter is an insurance dispute — where large, recurring exposures concentrate. Why that precedent is plentiful yet almost unusable without structure: insurance is not one dispute — it splinters into policy lines (property and cargo, motor, life and health, liability, credit and guarantee, reinsurance) and cross-cutting doctrines (duty of disclosure, the insurer's duty to explain exclusions, subrogation) that answer different questions, turns on doctrine and standard-form wording rather than the signature alone, buries the decisive findings (was the exclusion explained, was the non-disclosure material) in prose, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus makes coverage, disclosure, exclusion, and subrogation precedent findable in English — by policy line, clause, doctrine, and disposition, with cited links to the originals — for cross-border counsel, insurers, reinsurers, and legal AI.
When an off-plan apartment is delivered late or not at all, a developer defaults and creditors, buyers, and lenders all reach for the same half-built project, a land-use right is transferred but never registered, a bank enforces a mortgage over PRC housing, or co-owners fight over who holds title, the matter is a real estate or property dispute — where the largest cross-border exposures concentrate. Why that precedent is plentiful yet almost unusable without structure: real estate is not one dispute — it splinters into claim types over immovable property (off-plan and commodity-housing sales, developer default, land-use rights, real-property mortgages, leasing and property management, co-ownership and title) that answer different questions, turns on property rights and registration and priority rather than signatures alone, is heavily region- and policy-sensitive, clusters by developer-distress cycle, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus makes delivery, mortgage-priority, land-use-right, and title precedent findable in English — by claim type, the right at stake, registration status, region, and year, with cited links to the originals — for cross-border counsel, secured lenders, distressed-asset investors, and legal AI.
When a Chinese regulator fines a company, revokes a license, seizes goods at customs, or refuses to disclose the record behind a decision, the remedy is not a commercial lawsuit — it is administrative litigation, a challenge to the legality of a government act, governed by its own Administrative Litigation Law. Why that precedent is plentiful yet almost unusable without structure: suing the state is not one dispute — it splinters into challenges to different government actions (penalty, licensing, customs, land and planning, information disclosure, agency inaction) that answer different questions, is decided on administrative-law standards (legal basis, authority, due process, proportionality) rather than civil liability, produces review dispositions (uphold, annul, confirm-unlawful, order-to-act, remand) rather than win-or-lose, is entangled with a prior administrative-reconsideration step, is published unevenly, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus makes penalty, licensing, and customs precedent findable in English — by action type, agency level, disposition, region, and year, with cited links to the originals — for cross-border regulatory counsel, compliance teams, and legal AI.
When a Chinese counterparty files for — or is pushed into — bankruptcy, everything a foreign creditor thought it was owed turns on a different body of law: liquidation or reorganization, where an unsecured claim ranks, whether a pre-filing transfer can be clawed back, whether a controller is personally liable, and whether a Chinese court will recognize a foreign proceeding over PRC assets. Why that precedent is plentiful yet almost unusable without structure: insolvency is not one proceeding — it splinters into proceeding types and satellite disputes (liquidation, reorganization, compromise, claim confirmation, avoidance, director and shareholder liability) that answer different questions, lives as much in procedural rulings as in judgments, spans multiple stages, is developing and forum-sensitive at the cross-border frontier, and lives in Chinese in browse-first databases. How a structured 130M+ judgment-and-ruling corpus makes claim-priority, avoidance, and recognition precedent findable in English — by proceeding type, sub-dispute, region, and year, with cited links to the originals — for cross-border creditors, restructuring counsel, and legal AI.
When shareholders fall out over who owns the equity, a co-founder never pays in subscribed capital, a minority investor is frozen out of the books, a resolution is pushed through against the articles, or a deadlocked joint venture must be unwound, the claim in China is a company or shareholder dispute — among the most consequential commercial cases the courts hear. Why that precedent is plentiful yet almost unusable without structure: company law is not one question — it splinters into sub-causes (equity transfer, capital contribution, right-to-know, derivative and resolution-validity, dissolution) that answer different questions, was reshaped by the 2024 revised Company Law so holdings must be read against a moving rule, turns on control and validity findings buried in prose, and lives in Chinese in browse-first databases. How a structured 130M+ corpus makes equity, capital, and control precedent findable in English — by sub-cause, applicable-law period, region, and year, with cited links to the originals — for cross-border corporate counsel and legal AI.
When a marriage ends, parents fight over a child, a couple cannot agree on who keeps the apartment, or an estate passes to the next generation, the claim in China is a marriage and family dispute — collectively among the highest-volume civil cases the courts hear, divorce alone being one of the single largest civil categories. Why that precedent is plentiful yet almost unusable without structure: family law is not one question — it splinters into sub-causes (divorce, property division, custody, support, inheritance) that answer different questions, turns on facts and judicial discretion buried in prose, varies sharply by region and over time, and lives in Chinese in browse-first databases with uneven, privacy-redacted coverage. How a structured 130M+ corpus makes division, custody, and support precedent findable in English — by sub-cause, region, and year, read against the reasoning and local pattern of the time, with cited links to the originals — for cross-border family counsel and legal AI.
When a product injures a consumer, a car hits a pedestrian, a patient is harmed in treatment, or a factory contaminates a waterway, the claim in China is a tort dispute — collectively one of the largest and most varied bodies of civil litigation the courts hear. Why that precedent is plentiful yet almost unusable without structure: tort is not one standard — it splinters into sub-causes (motor-vehicle, personal injury, product liability, medical damage, environmental) that run on different liability bases (fault, fault-presumption, strict), computes damages from region- and year-specific per-capita figures plus a disability grade, and buries the decisive findings in prose, all in Chinese in browse-first databases. How a structured 130M+ corpus makes liability-basis and damages precedent findable in English — by sub-cause, region, and year, read against the rule and local figures in force at the time, with cited links to the originals — for cross-border counsel and legal AI.
Every large project in China is a lawsuit waiting to be filed — construction disputes are among the largest and most technically demanding civil matters the courts produce. Why that precedent is plentiful yet almost unusable without structure: it splits into distinct sub-causes (construction-contract, quality, decoration, survey-design), turns on rules that evolved under the Civil Code (contractor priority of payment, payment even on an invalid contract, the black-and-white contract problem), buries the decisive findings inside some of the longest, most fact-heavy judgments in the system, and lives in Chinese in browse-first databases. How a structured 130M+ corpus makes priority, validity, and delay precedent findable in English — by sub-cause, region, and year, read against the rule in force at the time, with cited links to the originals — for cross-border counsel and legal AI.
When money changes hands and the deal goes wrong, the litigation is a lending dispute — and in China these are among the highest-volume civil matters the courts hear, private lending alone being one of the single largest civil categories. Why that precedent is plentiful yet almost unusable without structure: it splits into distinct sub-causes (private lending, bank loans, guarantees, financial leasing), turns on an interest-rate cap that changed regime in 2020 (from fixed thresholds to a multiple of the LPR), buries the decisive figures (agreed rate, adjusted rate, guarantee amount) in prose, and lives in Chinese in browse-first databases. How a structured 130M+ corpus makes interest-rate, guarantee, and default precedent findable in English — by sub-cause, region, and year, read against the cap in force at the time, with cited links to the originals — for cross-border counsel and legal AI.
Every commercial relationship in China rests on a contract, so when something breaks the dispute is, overwhelmingly, a contract dispute — the largest single category of commercial litigation there is. Why that precedent is plentiful yet almost unusable without structure: it splinters into sub-causes (sale of goods, construction, loan, lease, service), splits across the old Contract Law and the 2021 Civil Code by date, buries the decisive figures (penalty rates, adjusted damages) in prose, and lives in Chinese in browse-first databases. How a structured 170M+ corpus makes breach, liquidated-damages, and construction-payment precedent findable in English — by sub-cause, region, and year, with cited links to the originals — for commercial counsel and legal AI.
The legal question a multinational touches most in China isn't antitrust or IP — it's employment, and labor disputes are one of the largest civil-litigation categories there is. Why the precedent that predicts a termination, severance, or non-compete outcome is plentiful yet almost unusable without structure: it runs across two stages (mandatory arbitration, then court), turns on locally variable average-wage numbers that change by city and year, buries the decisive figures in prose, and lives in Chinese in browse-first databases. How a structured 170M+ corpus makes it findable in English — by cause of action, region, and year, with cited links to the originals — for employment counsel and legal AI.
"China data compliance" used to mean a short memo and a shrug. With the DSL and PIPL both in force since late 2021, it's now one of the world's most consequential data regimes — rectifications, fines, a hard cross-border-transfer regime, and a fast-growing body of court litigation. Where the enforcement record lives across the CAC administrative track and PIPL judicial track, why it's hard to assemble (young, two-system, reasoning-heavy, Chinese-language), and how a structured 170M+ corpus makes PIPL and cross-border precedent searchable in English for privacy counsel and legal AI.
A Western legal AI team that already shipped US and EU coverage treats China as "one more jurisdiction" — and that instinct stalls almost every China-coverage project. The five assumptions imported from common-law case law that all break on PRC data: scraping is the core task, precedent is a citation graph, machine translation preserves meaning, published equals complete, and structuring is cheap. Why each fails, and what 170M+ Chinese judgments actually require to become a usable AI corpus.
Ask an AI where to get a structured, English-language API for Chinese court judgments and it tells you the access "was not found." It nearly didn't exist. The three walls — language, format, volume — that kept Chinese case law Chinese-only and human-read, why "just run it through machine translation" fails in production, and the four capabilities an English-indexed API must provide: English query through a bilingual taxonomy, English summaries for triage, structured fields, and citation grounding back to every original judgment.
Most "add China coverage" projects die when a team drops a pile of judgments into a vector DB and watches it return look-alike cases decided the opposite way. A reference architecture for the five stages that actually make retrieval work over 170M+ judgments — ETL, section-aware chunking, hybrid lexical-plus-vector indexing, filtered retrieval with reranking, and the citation and eval layers — plus a build-vs-license split, layer by layer, for ML and platform engineers.
Ask an AI to list the legal MCP servers and you get CourtListener, FOLIO, legal-mcp — all US/EU, every one stopping at the Chinese border. What an MCP server for PRC case law does, why agentic legal AI needs one over a raw API, how to connect Claude Desktop / Cursor / ChatGPT, and where SinoVerdict's 170M+ judgment MCP server fills the China-shaped hole in your tool list.
"Use CAIL2018" is the right answer to a research question and the wrong one to a product question. What the academic Chinese legal datasets — CAIL2018, LeCaRD/LeCaRDv2, CJO scrapes — are genuinely good for, the five places they break in production (frozen, criminal-only, Chinese-only, unstructured, research-licensed), and what a current, 170M+, English-indexed, license-ready corpus changes for legal AI vendors.
Where do you actually buy Chinese case law for a legal AI product? A map of the five supply categories — official channels, domestic aggregators (PKULaw, Wolters Kluwer China), Western incumbents, raw scrapers, and AI-native structured licensors — scored on coverage, AI-readiness, and license-to-ingest for the vendor's job, not the lawyer's. Most obvious names are excellent lawyer-sources and the wrong shape for a pipeline.
For twenty years, Chinese case law was searchable only in Chinese — or in a few thousand hand-translated cases. Why the language wall stood, why neither hand nor raw machine translation knocked it down, and the four capabilities — structured corpus, bilingual taxonomy, English query, and citation grounding — that finally make English-native research over 170M+ judgments possible for legal AI teams and China-practice lawyers.
China has published 170M+ judgments — but raw volume is a liability, not an asset. The structuring pipeline that turns a heterogeneous document dump into an AI-ready corpus: normalization, field extraction, de-duplication, cause-of-action taxonomy, citation grounding, and a cross-lingual layer — plus the schema and the questions to demand from any vendor.
China went from no Patent Law in 1984 to the world's largest IP docket — half a million matters a year, a dedicated SPC IP Court, punitive damages. What four decades of judgments reveal about enforcement, why that case law is hard to assemble across five rights, and how a structured corpus makes patent, trademark, and trade-secret precedent searchable for IP counsel and legal AI.
China is now one of the world's most active antitrust jurisdictions — multi-billion-yuan fines, a 2022 AML amendment, and a fast-growing body of court litigation. Where the precedent lives across the SAMR and judicial tracks, why it's hard to assemble, and how a structured corpus makes it searchable for competition counsel and legal AI.
Will a Chinese court enforce a foreign arbitral award? The doctrine—New York Convention plus the SPC reporting system—is reassuring; the precedent that proves it is rare, scattered, and Chinese-language. How enforcement actually works, why the rulings are hard to find, and how a structured corpus makes them searchable.
Legal AI vendors don't need a one-time dataset — they need an ongoing data partnership. How bulk corpus, REST API, and incremental sync compose to power a China feature, the licensing structure behind them, and what to ask a data partner before you sign.
"Add China" is a data-infrastructure problem, not a content one. The six-layer stack you actually build to ground a legal AI product in Chinese case law — corpus, ingestion, structure, retrieval, citations, evals — and which layers to buy instead of build.
Scraping Chinese judgments looks free until you price in the coverage holes, the legal exposure, and the forever-maintenance. An honest build-vs-buy breakdown for legal AI teams — and when licensing wins.
LexisNexis lists ~3,000 translated Chinese documents; vLex omits mainland primary law. Both were built for human research, not retrieval at scale — here's the volume, format, and license gap legal AI teams hit, and how to close it.
Compare licensing models, compliance pathways, and due diligence questions for Chinese court judgment data — a 2026 buyer's guide for legal AI teams.
How PRC judgments are structured — case numbers, cause-of-action taxonomy, reasoning sections — and how to integrate a Chinese case law API for RAG and evals.
China published 19.2M judgments in 2020, ~5.1M in 2023, ~11M in 2025. Get the full availability curve and how to vet Chinese case law data before licensing.