What Should a Chinese Legal Data Pilot Prove? Scoping, Exit Criteria and the Calendar You Cannot Compress
Almost every Chinese legal data pilot ends the same way. The team runs some queries, the results look reasonable, somebody builds a slide with three impressive examples and one honest caveat, and the meeting concludes that it looks promising and the next step is a bigger pilot.
That outcome is not a failure of the data or of the team. It is the predictable result of scoping a pilot to try the corpus rather than to decide something. A trial with no stated decision cannot fail, and a test that cannot fail carries no information — you spent a quarter to arrive back where you started, except the option now looks slightly better, because it was presented by people who wanted it to work.
This page is about the other kind of pilot: one scoped so that it can come back negative, and so that a negative result is as useful as a positive one. What follows is the decision rule to write before the data arrives, the five claims a pilot can genuinely test, the three it cannot, the calendar constraint no budget will move, and the clauses that belong in a pilot agreement rather than in the licence.
The failure mode has a name: no pre-registered decision
The distinction between an experiment and a demonstration is whether the interpretation was fixed before the observation. When retrieval quality is assessed after the fact by the people who chose the vendor, the assessment absorbs whatever came back: disappointing results become "a tuning problem", gaps become "out of scope for phase one". Nobody is being dishonest; this is simply what unstructured evaluation does.
The fix is procedural and costs nothing. Before the first record arrives, write down four things and circulate them to everyone who will be in the final meeting:
- The commitment waiting on this. Not "evaluate Chinese data" but the specific decision the pilot unblocks: ship a China module next roadmap cycle, add PRC coverage to an enterprise renewal, commit headcount to an ingestion pipeline. If nothing is waiting, you do not need a pilot yet; you need a scoping conversation.
- The claims being tested, each phrased so that it could turn out false.
- The exit criteria, with numbers you pick now and do not revise later.
- The named owner of the negative outcome — the person authorised to say "this did not clear the bar" without it reading as an attack on whoever sourced the vendor.
That last item is the one teams skip and the one that determines whether the rest matters.
PILOT DECISION MEMO (circulate before data arrives; do not revise after)
Decision this unblocks : ......................................
Decision date : ......................................
Owner of "no" : ......................................
Claim 1 .. 5 : (each one falsifiable; see table below)
Exit criteria : (numbers chosen now, by the people who will
have to live with them)
Out of scope : (explicitly: what this pilot will NOT settle)
Data disposition : (what happens to every delivered record on
the decision date, either way)
Five claims a pilot can actually test
A pilot is a measurement instrument with a narrow range, and these five are inside it. Each maps to a downstream decision — the test of whether it belongs in the pilot at all: if nothing changes based on the answer, do not spend the weeks.
| Claim | How it is measured | Decision it feeds |
|---|---|---|
| 1. The corpus contains what your users ask for | Known-item recall against a seed list you assembled yourself, before contact with the supplier | Whether the feature is buildable at all |
| 2. The fields your product reads are populated | Completeness per field, measured within each stratum you will query, never pooled across the whole sample | How much extraction and normalisation you must build |
| 3. Updates behave the way the roadmap assumes | Lag measured as a distribution across at least two deliveries separated by real weeks | Whether you can promise currency to your own customers |
| 4. The production path meets your budget | Latency and cost per operation on the delivery mode you will actually run, under a realistic access pattern | Architecture, and the price of the tier you need |
| 5. Your target task is answerable from these documents | A small adjudicated evaluation set, scored by someone qualified to read the judgments | Whether the product idea survives contact with the source material |
Claims 1 through 3 are coverage and quality questions, and the full protocol for running them is in the vendor verification piece. That is the instrument; this page is the experimental design around it.
Claim 5 is the one most often left out, the only one that tests your own product hypothesis rather than the supplier, and the one that most often comes back negative — because "the documents exist" and "the question is answerable from the documents" are different propositions. A judgment records what a court decided and the reasoning it chose to publish; it does not necessarily record the fact your feature needs. Finding that out during a pilot is cheap. Finding it out after you have built an ingestion pipeline is not — see the corpus preparation piece for what that pipeline involves once you commit to it.
Setting exit criteria you will honour
An exit criterion is only real if it was chosen by someone who will be inconvenienced by it. The mechanics that make this work:
- Set the bar per stratum, not overall. An aggregate number is dominated by whichever case types are most numerous, which are rarely the ones your customers care about. A corpus can look strong in aggregate and be unusable for the three case types your product is sold on.
- Decide the sample size before you see the first result, and be honest that a small stratum yields a wide interval. "We cannot tell from this sample" is a real finding, and a reason to extend the pilot rather than to round up.
- Write down what happens on a marginal result before the results exist. Marginal outcomes are where pre-registration earns its keep, because that is exactly when the arguments start.
Three claims a pilot cannot settle
Say these out loud at the start, because otherwise they will be quietly claimed in the write-up.
Long-run reliability. A supplier is at their most attentive during an evaluation. What you observe is their sales process, not their operations. Eight weeks of good behaviour tells you approximately nothing about the second year, and no amount of extending the pilot converts it into evidence. The instrument for this is contractual: service commitments, notice periods, remedies, and what happens when a delivery is late or defective. That is a term-sheet question, covered in the pricing and terms piece, not a pilot question.
Legal and compliance posture. How a sample performs tells you nothing about provenance, handling, or whether the deployment you intend is available to you. A pilot can surface the questions — the redaction specification, what survives in the records, which delivery modes exist — but the analysis runs in parallel, on its own timeline, and it is the one most likely to be the real critical path. Deployment geography is where that usually bites.
Production unit economics. Pilot volumes are unrepresentative in both directions: your engineers hammer the API in ways production never will, and production traffic has concurrency and caching behaviour the pilot never generated. Use the pilot for per-unit measurements, model the projection separately, and state the assumptions. Presenting a pilot bill as a forecast is how teams end up defending a number they never believed.
The calendar you cannot compress
This is the constraint that most pilot plans ignore, and it is the reason well-funded pilots still fail to answer the question they were commissioned for.
Claims 1, 2 and 5 are properties of a static delivery: given the sample and a competent engineer, they can be measured in days, and throwing more people at them genuinely does make them go faster.
Claim 3 is a property of a process, and it cannot be measured in days at any budget. To know how quickly new decisions appear, whether corrections propagate to records you already hold, and whether two deliveries are stable against each other, you need two deliveries separated by real weeks plus a re-pull of records you have seen before. Money does not buy this. The only way to have the answer by the decision date is to schedule the deliveries first and arrange everything else around them.
SCHEDULE BACKWARDS FROM THE DELIVERY CADENCE
decision date
^
| analysis + write-up (days)
| delivery #2 <-- re-pull of previously seen records
| ... real weeks in between, unavoidable ...
| delivery #1 <-- static tests (claims 1, 2, 5) start here
| seed list assembled BEFORE supplier contact
| legal/compliance review runs in PARALLEL from day one
start-|
Compressing the middle gap does not speed up the pilot.
It deletes claim 3 from the pilot, silently.
If your decision date genuinely does not allow two separated deliveries, that is a legitimate position — but then drop claim 3 explicitly, record that update behaviour is being taken on contractual commitment rather than measurement, and make sure the agreement carries a remedy for it. What you must not do is run a compressed pilot and let the write-up imply that freshness was verified.
Run it on the topology you will actually operate
Pilot findings do not transfer cleanly across delivery modes, and this is where a lot of pilot effort is wasted. A bulk sample in a local index says nothing reliable about per-query latency across a cross-border link; an API-based pilot says nothing about whether your training pipeline can chew through full text at volume. Measuring one and deciding the other is the most common way a pilot produces a confidently wrong number.
The delivery modes — bulk drop, incremental sync, query API, tool-server access — differ in what they can even demonstrate, as laid out in the partnership models piece. Where the production topology is still open, the pragmatic order is to pilot the query-only path first: it stands up fastest, it is reversible, and it forces the integration questions early. Then write the right to migrate to another delivery mode, on the same commercial terms, into the pilot agreement — so the fast path now does not silently decide the architecture later.
Whichever mode you pilot, wire it against the real field structure rather than a hand-cleaned extract: a team that spends the pilot writing parsers is measuring its own parser, not the corpus. The API structure walkthrough and the retrieval pipeline piece cover that integration surface.
Seven clauses that belong in the pilot agreement
These are cheap to agree beforehand and expensive to argue afterwards. They get missed because they are distinct from the licence terms: the commercial team is focused on the deal that follows, and the pilot paper is treated as a formality.
PILOT AGREEMENT — the seven that are not in the licence
1. DATA DISPOSITION What happens to every delivered record on the
decision date, either way, and how deletion is
evidenced. Agree this while you still have leverage.
2. FEE CREDIT Whether pilot fees credit against a subsequent
licence, in full or in part, and for how long
that credit stays available.
3. MIGRATION RIGHT The right to move to a different delivery mode
on the same commercial terms — so the fast pilot
path does not lock the production architecture.
4. ARTEFACT OWNERSHIP Your evaluation sets, seed lists, field mappings
and annotations are yours. They are usually the
most valuable thing the pilot produces.
5. NO AUTO-CONVERSION The pilot does not roll into a paid term by
default, and silence is not acceptance.
6. EXTENSION MECHANISM A defined way to add weeks. Pilots gated on
calendar time need this more often than not.
7. PUBLICITY SCOPE Neither side describes the engagement externally
in terms the other has not agreed in writing.
Clause 4 is the one worth arguing for. By the end of a serious pilot your team has built a seed list, an adjudicated evaluation set and a field mapping specific to your task — work that is reusable against the next supplier, the next corpus and the next model, and the strongest thing you can carry into a renegotiation. Do not let it become a joint work product by inattention.
Four ways a pilot misleads you
Even a well-designed pilot has characteristic failure modes. Check for these before writing up:
- Sample selection you did not control. If the supplier chose which records you received, you measured their selection, not their corpus. Specify the sample yourself — by case type, court level, date range and region — and include strata you expect to be weak.
- Seed list contamination. A seed list assembled after conversations with the supplier drifts toward what came up in them. Build it first, from your own matters or users' queries, and store it before the first call.
- Aggregate success hiding a fatal stratum. Strong overall numbers with one weak case type is a pass on paper and a failure in market if that case type is what you sell. This is why exit criteria are set per stratum.
- The demo effect. The queries run in front of the supplier are not a random sample of your workload. Reserve a set of questions that nobody outside your team has seen, and run them last.
What this page does not settle
A well-scoped pilot narrows uncertainty; it does not eliminate it. It will not tell you whether the supplier will still be operating well in two years, whether your intended deployment is available to you, or what the arrangement costs at production scale. Nor does it substitute for the sourcing decision upstream of it — whether to license at all, and on what basis — covered in the licensing guide and, on building it yourself, in the piece on what the source portal's robots.txt actually says.
What a pilot can do is convert an argument about opinions into an argument about measurements, and give a team permission to say no on evidence. That is a smaller claim than most pilot charters make, and it is worth the quarter.
Frequently asked questions
One decision, stated before the data arrives. A pilot is worth running when there is a specific commitment waiting on it — whether to build a China module into the roadmap, for instance — and when you can name in advance the observation that would cause you not to make that commitment. The testable claims are narrow: that the corpus retrieves what your own question set requires, that the fields your product reads are populated within the strata you actually query, that freshness behaves as a distribution rather than a best case, that the delivery path you will operate meets your latency and cost budget, and that your target task is answerable from these documents at all. Anything broader is a demonstration, not an experiment.
Long enough for the properties that only reveal themselves over calendar time. Retrieval quality and field completeness are properties of a static delivery and can be measured in days. Freshness lag, delivery stability and whether corrections propagate cannot be measured in days at any budget: they require at least two deliveries separated by real weeks plus a re-pull of records you have already seen. If your pilot must answer questions about update behaviour, schedule the deliveries first and fit the analysis around them.
Three things worth naming so nobody claims otherwise in the write-up. Long-run reliability: an attentive evaluation period tells you about the sales process, not the operations — that belongs in the service commitments of the agreement. The legal and compliance analysis, which turns on provenance and handling rather than sample performance. And production unit economics, because pilot volumes and access patterns are unrepresentative of steady state.
On the delivery mode you intend to operate in production, because findings do not transfer cleanly between them. A bulk sample in a local index says nothing reliable about per-query latency over a cross-border link, and an API pilot says nothing about whether your training pipeline handles full text at volume. Where the production topology is still open, pilot the query-only path first — it is fastest to stand up and reversible — and write the right to migrate to another delivery mode into the pilot agreement so the architecture stays open.
Data disposition at the end of the pilot and how deletion is evidenced; whether pilot fees credit against a subsequent licence; the right to migrate to a different delivery mode on the same commercial terms; ownership of the artefacts your team creates, such as evaluation sets and field mappings; the absence of automatic conversion to a paid term; a defined extension mechanism, since calendar-gated pilots often need one; and the scope of publicity on both sides.
Run the pilot against a real delivery, not a curated sample.
Tell us the decision the pilot has to unblock and the strata your product actually queries, and we will scope a time-boxed evaluation against the 160M+ record corpus on the delivery mode you intend to run — with the sample specified by you, the field structure documented up front, and a delivery cadence that lets you measure update behaviour rather than take it on assertion. A trial API key comes with it. Write to chenjiaxin@wenshucha.com or use the form.
Request trial access