Buyer's protocol

What Should a Chinese Legal Data Pilot Prove? Scoping, Exit Criteria and the Calendar You Cannot Compress

Almost every Chinese legal data pilot ends the same way. The team runs some queries, the results look reasonable, somebody builds a slide with three impressive examples and one honest caveat, and the meeting concludes that it looks promising and the next step is a bigger pilot.

That outcome is not a failure of the data or of the team. It is the predictable result of scoping a pilot to try the corpus rather than to decide something. A trial with no stated decision cannot fail, and a test that cannot fail carries no information — you spent a quarter to arrive back where you started, except the option now looks slightly better, because it was presented by people who wanted it to work.

This page is about the other kind of pilot: one scoped so that it can come back negative, and so that a negative result is as useful as a positive one. What follows is the decision rule to write before the data arrives, the five claims a pilot can genuinely test, the three it cannot, the calendar constraint no budget will move, and the clauses that belong in a pilot agreement rather than in the licence.

Commercial and technical guidance for buyers evaluating a Chinese case law corpus. Not legal advice. Thresholds and figures in the templates below are placeholders you are meant to replace with your own; the point of the exercise is that you choose the numbers before you see any results.

The failure mode has a name: no pre-registered decision

The distinction between an experiment and a demonstration is whether the interpretation was fixed before the observation. When retrieval quality is assessed after the fact by the people who chose the vendor, the assessment absorbs whatever came back: disappointing results become "a tuning problem", gaps become "out of scope for phase one". Nobody is being dishonest; this is simply what unstructured evaluation does.

The fix is procedural and costs nothing. Before the first record arrives, write down four things and circulate them to everyone who will be in the final meeting:

  1. The commitment waiting on this. Not "evaluate Chinese data" but the specific decision the pilot unblocks: ship a China module next roadmap cycle, add PRC coverage to an enterprise renewal, commit headcount to an ingestion pipeline. If nothing is waiting, you do not need a pilot yet; you need a scoping conversation.
  2. The claims being tested, each phrased so that it could turn out false.
  3. The exit criteria, with numbers you pick now and do not revise later.
  4. The named owner of the negative outcome — the person authorised to say "this did not clear the bar" without it reading as an attack on whoever sourced the vendor.

That last item is the one teams skip and the one that determines whether the rest matters.

PILOT DECISION MEMO  (circulate before data arrives; do not revise after)

Decision this unblocks : ......................................
Decision date          : ......................................
Owner of "no"          : ......................................

Claim 1 .. 5           : (each one falsifiable; see table below)
Exit criteria          : (numbers chosen now, by the people who will
                          have to live with them)
Out of scope           : (explicitly: what this pilot will NOT settle)
Data disposition       : (what happens to every delivered record on
                          the decision date, either way)

Five claims a pilot can actually test

A pilot is a measurement instrument with a narrow range, and these five are inside it. Each maps to a downstream decision — the test of whether it belongs in the pilot at all: if nothing changes based on the answer, do not spend the weeks.

ClaimHow it is measuredDecision it feeds
1. The corpus contains what your users ask forKnown-item recall against a seed list you assembled yourself, before contact with the supplierWhether the feature is buildable at all
2. The fields your product reads are populatedCompleteness per field, measured within each stratum you will query, never pooled across the whole sampleHow much extraction and normalisation you must build
3. Updates behave the way the roadmap assumesLag measured as a distribution across at least two deliveries separated by real weeksWhether you can promise currency to your own customers
4. The production path meets your budgetLatency and cost per operation on the delivery mode you will actually run, under a realistic access patternArchitecture, and the price of the tier you need
5. Your target task is answerable from these documentsA small adjudicated evaluation set, scored by someone qualified to read the judgmentsWhether the product idea survives contact with the source material

Claims 1 through 3 are coverage and quality questions, and the full protocol for running them is in the vendor verification piece. That is the instrument; this page is the experimental design around it.

Claim 5 is the one most often left out, the only one that tests your own product hypothesis rather than the supplier, and the one that most often comes back negative — because "the documents exist" and "the question is answerable from the documents" are different propositions. A judgment records what a court decided and the reasoning it chose to publish; it does not necessarily record the fact your feature needs. Finding that out during a pilot is cheap. Finding it out after you have built an ingestion pipeline is not — see the corpus preparation piece for what that pipeline involves once you commit to it.

Setting exit criteria you will honour

An exit criterion is only real if it was chosen by someone who will be inconvenienced by it. The mechanics that make this work:

Three claims a pilot cannot settle

Say these out loud at the start, because otherwise they will be quietly claimed in the write-up.

Long-run reliability. A supplier is at their most attentive during an evaluation. What you observe is their sales process, not their operations. Eight weeks of good behaviour tells you approximately nothing about the second year, and no amount of extending the pilot converts it into evidence. The instrument for this is contractual: service commitments, notice periods, remedies, and what happens when a delivery is late or defective. That is a term-sheet question, covered in the pricing and terms piece, not a pilot question.

Legal and compliance posture. How a sample performs tells you nothing about provenance, handling, or whether the deployment you intend is available to you. A pilot can surface the questions — the redaction specification, what survives in the records, which delivery modes exist — but the analysis runs in parallel, on its own timeline, and it is the one most likely to be the real critical path. Deployment geography is where that usually bites.

Production unit economics. Pilot volumes are unrepresentative in both directions: your engineers hammer the API in ways production never will, and production traffic has concurrency and caching behaviour the pilot never generated. Use the pilot for per-unit measurements, model the projection separately, and state the assumptions. Presenting a pilot bill as a forecast is how teams end up defending a number they never believed.

The calendar you cannot compress

This is the constraint that most pilot plans ignore, and it is the reason well-funded pilots still fail to answer the question they were commissioned for.

Claims 1, 2 and 5 are properties of a static delivery: given the sample and a competent engineer, they can be measured in days, and throwing more people at them genuinely does make them go faster.

Claim 3 is a property of a process, and it cannot be measured in days at any budget. To know how quickly new decisions appear, whether corrections propagate to records you already hold, and whether two deliveries are stable against each other, you need two deliveries separated by real weeks plus a re-pull of records you have seen before. Money does not buy this. The only way to have the answer by the decision date is to schedule the deliveries first and arrange everything else around them.

SCHEDULE BACKWARDS FROM THE DELIVERY CADENCE

  decision date
        ^
        |  analysis + write-up            (days)
        |  delivery #2  <-- re-pull of previously seen records
        |  ... real weeks in between, unavoidable ...
        |  delivery #1  <-- static tests (claims 1, 2, 5) start here
        |  seed list assembled BEFORE supplier contact
        |  legal/compliance review runs in PARALLEL from day one
  start-|

Compressing the middle gap does not speed up the pilot.
It deletes claim 3 from the pilot, silently.

If your decision date genuinely does not allow two separated deliveries, that is a legitimate position — but then drop claim 3 explicitly, record that update behaviour is being taken on contractual commitment rather than measurement, and make sure the agreement carries a remedy for it. What you must not do is run a compressed pilot and let the write-up imply that freshness was verified.

Run it on the topology you will actually operate

Pilot findings do not transfer cleanly across delivery modes, and this is where a lot of pilot effort is wasted. A bulk sample in a local index says nothing reliable about per-query latency across a cross-border link; an API-based pilot says nothing about whether your training pipeline can chew through full text at volume. Measuring one and deciding the other is the most common way a pilot produces a confidently wrong number.

The delivery modes — bulk drop, incremental sync, query API, tool-server access — differ in what they can even demonstrate, as laid out in the partnership models piece. Where the production topology is still open, the pragmatic order is to pilot the query-only path first: it stands up fastest, it is reversible, and it forces the integration questions early. Then write the right to migrate to another delivery mode, on the same commercial terms, into the pilot agreement — so the fast path now does not silently decide the architecture later.

Whichever mode you pilot, wire it against the real field structure rather than a hand-cleaned extract: a team that spends the pilot writing parsers is measuring its own parser, not the corpus. The API structure walkthrough and the retrieval pipeline piece cover that integration surface.

Seven clauses that belong in the pilot agreement

These are cheap to agree beforehand and expensive to argue afterwards. They get missed because they are distinct from the licence terms: the commercial team is focused on the deal that follows, and the pilot paper is treated as a formality.

PILOT AGREEMENT — the seven that are not in the licence

1. DATA DISPOSITION      What happens to every delivered record on the
                         decision date, either way, and how deletion is
                         evidenced. Agree this while you still have leverage.

2. FEE CREDIT            Whether pilot fees credit against a subsequent
                         licence, in full or in part, and for how long
                         that credit stays available.

3. MIGRATION RIGHT       The right to move to a different delivery mode
                         on the same commercial terms — so the fast pilot
                         path does not lock the production architecture.

4. ARTEFACT OWNERSHIP    Your evaluation sets, seed lists, field mappings
                         and annotations are yours. They are usually the
                         most valuable thing the pilot produces.

5. NO AUTO-CONVERSION    The pilot does not roll into a paid term by
                         default, and silence is not acceptance.

6. EXTENSION MECHANISM   A defined way to add weeks. Pilots gated on
                         calendar time need this more often than not.

7. PUBLICITY SCOPE       Neither side describes the engagement externally
                         in terms the other has not agreed in writing.

Clause 4 is the one worth arguing for. By the end of a serious pilot your team has built a seed list, an adjudicated evaluation set and a field mapping specific to your task — work that is reusable against the next supplier, the next corpus and the next model, and the strongest thing you can carry into a renegotiation. Do not let it become a joint work product by inattention.

Four ways a pilot misleads you

Even a well-designed pilot has characteristic failure modes. Check for these before writing up:

What this page does not settle

A well-scoped pilot narrows uncertainty; it does not eliminate it. It will not tell you whether the supplier will still be operating well in two years, whether your intended deployment is available to you, or what the arrangement costs at production scale. Nor does it substitute for the sourcing decision upstream of it — whether to license at all, and on what basis — covered in the licensing guide and, on building it yourself, in the piece on what the source portal's robots.txt actually says.

What a pilot can do is convert an argument about opinions into an argument about measurements, and give a team permission to say no on evidence. That is a smaller claim than most pilot charters make, and it is worth the quarter.

Frequently asked questions

What should a Chinese legal data pilot be scoped to prove?

One decision, stated before the data arrives. A pilot is worth running when there is a specific commitment waiting on it — whether to build a China module into the roadmap, for instance — and when you can name in advance the observation that would cause you not to make that commitment. The testable claims are narrow: that the corpus retrieves what your own question set requires, that the fields your product reads are populated within the strata you actually query, that freshness behaves as a distribution rather than a best case, that the delivery path you will operate meets your latency and cost budget, and that your target task is answerable from these documents at all. Anything broader is a demonstration, not an experiment.

How long should a legal data pilot run?

Long enough for the properties that only reveal themselves over calendar time. Retrieval quality and field completeness are properties of a static delivery and can be measured in days. Freshness lag, delivery stability and whether corrections propagate cannot be measured in days at any budget: they require at least two deliveries separated by real weeks plus a re-pull of records you have already seen. If your pilot must answer questions about update behaviour, schedule the deliveries first and fit the analysis around them.

What can a pilot not tell you about a supplier?

Three things worth naming so nobody claims otherwise in the write-up. Long-run reliability: an attentive evaluation period tells you about the sales process, not the operations — that belongs in the service commitments of the agreement. The legal and compliance analysis, which turns on provenance and handling rather than sample performance. And production unit economics, because pilot volumes and access patterns are unrepresentative of steady state.

Should a data pilot run on bulk delivery or on the API?

On the delivery mode you intend to operate in production, because findings do not transfer cleanly between them. A bulk sample in a local index says nothing reliable about per-query latency over a cross-border link, and an API pilot says nothing about whether your training pipeline handles full text at volume. Where the production topology is still open, pilot the query-only path first — it is fastest to stand up and reversible — and write the right to migrate to another delivery mode into the pilot agreement so the architecture stays open.

What should be in a pilot agreement that is not in the licence?

Data disposition at the end of the pilot and how deletion is evidenced; whether pilot fees credit against a subsequent licence; the right to migrate to a different delivery mode on the same commercial terms; ownership of the artefacts your team creates, such as evaluation sets and field mappings; the absence of automatic conversion to a paid term; a defined extension mechanism, since calendar-gated pilots often need one; and the scope of publicity on both sides.

Run the pilot against a real delivery, not a curated sample.

Tell us the decision the pilot has to unblock and the strata your product actually queries, and we will scope a time-boxed evaluation against the 160M+ record corpus on the delivery mode you intend to run — with the sample specified by you, the field structure documented up front, and a delivery cadence that lets you measure update behaviour rather than take it on assertion. A trial API key comes with it. Write to chenjiaxin@wenshucha.com or use the form.

Request trial access