Commercial terms

What Does Licensing Chinese Court Data Cost? Pricing Structures and the Term Sheet That Moves the Number

The question arrives in almost every first email, usually in the last line: what does this cost? It is the right question and it has no honest one-number answer, which is why most vendor pages avoid it entirely and most buyers leave the call knowing less than when they joined.

So here is the useful version instead. This piece sets out the five variables that determine the number, the five pricing structures actually used in data licensing and who each one favours, and the nine contract clauses that decide whether the corpus you paid for is usable in the product you are building. It ends with a one-page scoping spec you can send to any supplier — us included — to get a real quote in one round rather than six.

Commercial guidance only, not legal advice. Contract terms are jurisdiction- and deal-specific; have your own counsel review anything you sign. No prices are quoted here, for the reason set out immediately below.

Why nobody publishes a price, and what actually follows from that

Data licensing is priced per deal because the same corpus is several different products depending on what you are permitted to do with it. A frozen snapshot for an academic benchmark and a production dependency carrying training rights, redistribution to your own enterprise customers and a refresh obligation are not the same transaction with a different discount — they carry different obligations for the licensor over different time horizons.

Two things follow, and only two. First, any article quoting you a dollar-per-record figure for PRC case law is guessing, and you should treat the number the way you would treat a valuation from someone who has not seen the accounts. Second — and this is the actionable half — the buyer who arrives with scope already defined gets a faster and usually lower quote, because an undefined scope has to be priced for its worst plausible reading. If a supplier cannot tell whether you intend to train foundation models on the corpus and resell extracts to two hundred law firms, the quote has to assume you might.

The five variables that set the number

Every quote in this category is a function of the same five inputs. Fix them before you ask.

VariableWhat moves the priceWhere teams under-specify
1. ScopeHow much of the corpus: all case types or a slice, which years, which court levels, which regionsAsking for "everything" out of caution when the product only ever queries three case categories
2. Field depthMetadata only (case number, court, date, cause of action, parties, disposition) versus full reasoning text versus normalised derived fieldsAssuming full text is included by default, then discovering the index cannot quote the passage the model cited
3. Delivery modeOne-time bulk transfer, hosted REST API, MCP access, incremental sync, or a combinationBuying bulk to save money, then rebuilding a retrieval service that was available as a line item
4. Rights grantedInternal research, retrieval indexing, embedding generation, fine-tuning, pre-training, display to end users, redistributionTreating these as one permission called "use" — see the term sheet below
5. Term and refreshLength of licence, refresh cadence, whether new documents during the term are included or priced separatelySigning a perpetual snapshot for a product whose whole value proposition is currency

Notice that only the first two are about data volume. The last three are about risk allocation, and in most negotiations they move the number further than scope does.

Five pricing structures, and who each one favours

These are the structures used across commercial data licensing generally; the Chinese case law market uses the same vocabulary. The right choice is less about which is cheapest on day one than about which risk you would rather carry.

StructureFitsCarries the riskWatch for
One-time bulk feeFine-tuning, benchmarking, a frozen research snapshotYou carry staleness riskWhat "delivery complete" means, and whether re-delivery after a pipeline failure is chargeable
Annual subscription with refreshAny product answering questions about current lawLicensor carries staleness riskWhether the refresh is contractual or best-efforts, and what the cadence actually is
Metered APIPilots, spiky workloads, unproven demandYou carry volume riskCost scaling with your success; agree a ceiling or a conversion trigger up front
Per-seat downstreamVendors reselling into a defined user baseShared, roughly tracks your revenueReporting overhead, and the definition of a "seat" when your users are agents rather than people
Revenue shareEarly-stage buyers with no budgetLicensor carries adoption riskThe cheapest structure if the product fails and usually the most expensive if it works

One pattern worth naming because it recurs: teams pick metered API to keep the pilot cheap, ship successfully, and discover at scale that a subscription would have cost less than the metered bill for eighteen months. If your roadmap says the workload grows, negotiate the conversion path into the first agreement rather than renegotiating from a position where switching costs are already sunk.

The term sheet: nine clauses that decide whether it is usable

Price is what you argue about. These are what determine whether the thing you bought works in production. Ordered roughly by how expensive they are to get wrong.

1. Grant scope, stated as named uses

Ask for each use to be listed explicitly: pre-training, fine-tuning, evaluation, embedding generation, retrieval indexing, display of extracts to end users, and inclusion of extracts in generated output. A licence drafted for a search product may permit indexing and display while saying nothing about model weights. Silence is not permission, and the party who eventually asks you to evidence the grant is usually not the licensor but your own customer's procurement team.

2. Survival on termination

The most expensive clause in the document and rarely the most negotiated. What happens, when the agreement ends, to a model already fine-tuned on the corpus, to an embedding index, to a derived citation graph, to documents cached inside a customer's deployment? If the answer is "delete all derivatives", then a model trained during the term becomes contractually unusable on the day the relationship ends, and your switching cost is not the next licence fee — it is a retraining programme plus whatever you promised your customers. Settle it in writing before signature, while leverage is still symmetrical.

3. Redistribution and sublicensing

Can you show extracts to your own customers? Can they export them? Can a customer's on-premises deployment hold a copy? Legal AI products almost always redistribute in some form, and a licence that permits internal use only is a product blocker discovered at the worst possible moment, typically during a large customer's security review.

4. Coverage and completeness commitments

Distinguish clearly between what the vendor holds and what exists. A commitment worth having describes the delivered corpus — counts by year, court level, region and case type — rather than promising the universe of PRC judgments. Ask for that breakdown as a diligence artefact and test retrieval against ten matters you already know before signing. This is also the honest limit of our own claims: we describe a corpus of more than 160 million publicly available PRC court judgment records, and coverage questions are answered with a breakdown rather than an adjective.

5. Refresh cadence and the definition of "current"

"Regularly updated" is not a term. Ask for cadence, expected lag between a document becoming publicly available and appearing in your delivery, and what happens when an upstream source changes its publication behaviour. The last one matters more in this jurisdiction than in most, because the supply side has not been static.

6. Takedown and deletion propagation

Publicly available court documents are sometimes withdrawn or corrected at source, and parties do make removal requests. Your contract should say who is obliged to act, how the instruction reaches you, and whether you must propagate deletion into derived indexes and customer deployments. If a vendor has no process to describe here, they have not run one.

7. Compliance representations and transfer pathway

Where will the data sit, who operates the transfer, and which representations does each side make? Cross-border data movement out of the PRC is a regulated area and the answer belongs to counsel, not to a blog post — but the allocation of that work between the parties is a commercial term you negotiate, and leaving it unstated does not make it go away.

8. Audit rights, both directions

Usage audits are standard. Read them for scope and notice, and ask for the reciprocal: your right to verify what you were delivered against what was described. Most drafts give the licensor inspection rights and the licensee none.

9. Assignment and change of control

If either party is acquired, does the licence follow? Data suppliers get bought, and a supply agreement that terminates on change of control is a dependency your own acquirer will find in diligence.

The scoping spec that gets you a real quote

Send this instead of booking a discovery call. Suppliers can price against it directly, and you will find that writing it settles internal arguments you did not know you were having.

1. Scope        case types / year range / court levels / regions
2. Field depth  metadata only | + full reasoning text | + normalised derived fields
3. Delivery     bulk | REST API | MCP | incremental sync | combination
4. Rights       name each: pre-train, fine-tune, eval, embeddings,
                retrieval index, display extracts, extracts in output
5. Term         length + refresh cadence + are new documents included
6. Redistribute to your customers? in what form? on-prem copies?
7. Deployment   where the data will be hosted and processed
8. Timeline     target start date, pilot vs production

Eight lines. In our experience the deals that close quickly are the ones where the buyer sent something like this in the first message, and the deals that stall for a quarter are the ones where item 4 was never resolved internally.

Four red flags in a data quote

What this does not tell you

Four honest limits, because a page about commercial terms that pretends to be complete is the least trustworthy kind.

  1. No numbers. We do not publish a price list, and this page deliberately quotes none — not as a negotiating posture but because a figure detached from scope would mislead every reader whose scope differs.
  2. Not legal advice, and specifically not on cross-border transfer. Clause 7 above is where the real legal work sits, and it belongs with qualified counsel in your jurisdictions.
  3. Structures, not market rates. The five pricing models are general commercial data licensing practice. We are not claiming to know what any other supplier charges.
  4. Records, not full-text guarantees. We describe a corpus of more than 160 million publicly available PRC court judgment records. Field completeness varies across decades of source formatting, which is exactly why clause 4 asks for a breakdown rather than an adjective — including from us.

If you want the surrounding context before you scope anything, the 2026 licensing guide covers the supply picture and diligence questions, the partnership models piece works through bulk, API and sync as delivery modes, and license vs scrape deals with the build-versus-buy decision that usually precedes this one. On the build side specifically, what the source portal's robots.txt actually says is worth ten minutes before anyone budgets a crawler. For integration shape, see the API structure walkthrough and the MCP server.

Frequently asked questions

How much does it cost to license Chinese court judgment data?

There is no list price for this category, and any figure quoted without your scope attached is a guess. Data licensing is priced per deal because the same corpus can be sold as a one-time research snapshot or as a production dependency with training rights, redistribution to your own customers and a refresh obligation, and those are different products with different risk profiles for the licensor. What you can do before contacting anyone is fix the five variables that set the number: how much of the corpus you need, how deep the fields go, how it is delivered, which rights are granted, and how long the term runs with what refresh cadence. A vendor who receives those five answers can usually price in one round instead of six.

Which pricing structure is best for a legal AI company?

It depends on which risk you would rather carry. A one-time bulk fee gives you a fixed, forecastable cost and leaves you carrying the staleness risk, which suits a fine-tuning or benchmarking programme against a frozen snapshot. An annual subscription with refresh moves staleness to the licensor and suits any product that answers questions about current law. Metered API pricing is the cheapest way to start and the easiest way to be surprised later, because cost scales with your success. Per-seat downstream pricing aligns cost with revenue but requires you to report seats. Revenue share sounds attractive when you have no budget and tends to be the most expensive structure if the product works.

What is the most expensive clause to get wrong in a data licence?

Survival on termination, and it is rarely the clause teams negotiate hardest. The question is what happens to artefacts you have already created when the agreement ends: a fine-tuned model, an embedding index, a derived citation graph, cached documents in a customer deployment. If the contract requires deletion of all derivatives on termination, then a model you trained during the term becomes contractually unusable the day the relationship ends, and the switching cost is not the next licence fee but a retraining programme plus whatever your customers were promised. Settle in writing which artefacts survive, for how long, and in what form before you sign, not at renewal when leverage has moved.

Does a data licence let me train a model on the corpus?

Only if it says so. Treat training, retrieval-augmented indexing and embedding generation as three separate permissions rather than one, because a licence written for search products may permit indexing and display while saying nothing about model weights. Ask for each use to be named explicitly in the grant: pre-training, fine-tuning, evaluation, embedding generation, retrieval indexing, display of extracts to end users, and inclusion of extracts in generated output. Silence is not permission, and the party who will eventually ask you to prove the grant is not the licensor but your own enterprise customer's procurement team.

What should I send a vendor to get a real quote quickly?

A one-page scoping spec beats a discovery call. State the case-type and year scope you need; the field depth, distinguishing metadata-only from full reasoning text; the delivery mode, whether bulk, API, incremental sync or a combination; the exact rights you need, using the seven-item list of named uses; the term and refresh cadence; whether you will redistribute extracts to your own customers and in what form; your deployment geography; and your target start date. Vendors quote faster and lower against a defined scope, because an undefined scope has to be priced for its worst plausible reading.

Why is Chinese court data priced differently from other jurisdictions?

Because the supply side is different, not because the market is opaque for its own sake. In jurisdictions with a stable bulk feed, competing vendors price against the same underlying availability, so quotes converge. For PRC judgments, what any given vendor holds depends on when and how it was collected, so coverage, field normalisation and refresh capability vary materially between suppliers and are the substance of what you are buying. That is why coverage verification belongs in diligence rather than after signature: ask for a breakdown by year, court level, region and case type, and test retrieval against cases you already know before you commit.

Bring the eight-line spec. We will answer against it.

Send the scoping spec above and you will get a scoped answer rather than a brochure — including a coverage breakdown of the 160M+ record corpus by year, court level, region and case type, and a trial API key so your team can test retrieval against matters you already know before any commercial discussion. Write to chenjiaxin@wenshucha.com or use the form.

Request trial access