Buyer's question

The Year China’s Bankruptcy Registry Outgrew Its Own Pagination Window

China's national enterprise bankruptcy registry told us this morning that it holds 908,378 announcements. It will serve us 500 of them. That ratio — 0.055 per cent — is the same one we published yesterday, and it is not the interesting part.

The interesting part is that this registry was fully walkable, one calendar day at a time, until about three years ago. Nothing was switched off. The window did not shrink. The publication rate grew past it. A pipeline written in 2021 that sliced by publication date and paged each day to exhaustion would have been complete when it shipped, would have stayed complete through 2022, and would have begun dropping roughly a third of every working day somewhere in 2023 — returning HTTP 200 the entire time.

We also have to correct ourselves. Yesterday we reported that 8 June 2025 was an empty Sunday on this chain and that the announcement list could not be rescued by filtering. Both of those are wrong, and today's measurements are checkable in a way yesterday's were not.

Scope, stated once. Roughly 330 HTTP requests from one machine on 2 September 2026, issued serially with 3.2 seconds between them, all against one host. Everything below is from served markup — what a program receives — not from a rendered browser view. robots.txt was requested before any list endpoint and returned HTTP 404, as it did yesterday. Nothing here is a legal conclusion about what the site permits: we did not log in, did not read terms of service, and did not deliberately probe for rate limits. No personal data from any record is reproduced. Every figure is dated, because this behaviour can change without notice.

A better instrument than yesterday's

Yesterday we could only quote the page-count literal that the site's own pagination widget is initialised with, and we said plainly that we had no way to check it. Today we found the second number. Every filtered response carries a record count in its markup — the exact size of the result set, not a page estimate — and that changes what can be claimed.

It also let us validate the instrument instead of trusting it. On three days small enough to walk end to end, we compared the advertised count against what the pages actually delivered:

Day sliceCount advertisedPages advertisedDistinct records walkedDuplicatesResponse past the last page
8 June 2025465460HTTP 200, ~3,418 bytes, 0 rows
2 June 2025606600HTTP 200, ~3,418 bytes, 0 rows
15 June 2025105111050HTTP 200, ~3,420 bytes, 0 rows

Three for three, exactly, with no record served twice. The date filter is genuinely filtering, too: every row returned for a given day carried that day's publication date in its own markup, and a range covering the whole of 1900 returned zero rows rather than the unfiltered list. So when this registry states the size of a slice, it is telling the truth — at least for slices small enough that the truth can be checked, which is a caveat worth holding onto, because those are exactly the slices you do not need help with.

This is also how we know yesterday's reading of 8 June was wrong. That Sunday is not empty. It holds 46 announcements, the count header says 46, a full five-page walk yields 46 distinct identifiers, and each of them displays 8 June 2025. We have not reconstructed which parameter yesterday's run got wrong; we can only report that today's number is corroborated three ways and yesterday's was not.

Nine days out of fourteen do not fit

Yesterday's other claim was that a single busy day could overflow the 500-record window, based on one day. That day was not an outlier. It is the norm, and the exceptions are weekends.

DayWeekdayAnnouncementsFits in the 500-record window?
2 June 2025Mon60Yes — public holiday
3 June 2025Tue719No
4 June 2025Wed788No
5 June 2025Thu855No
6 June 2025Fri820No
7 June 2025Sat127Yes
8 June 2025Sun46Yes
9 June 2025Mon847No
10 June 2025Tue821No
11 June 2025Wed763No
12 June 2025Thu783No
13 June 2025Fri889No
14 June 2025Sat112Yes
15 June 2025Sun105Yes

Every ordinary working day in that fortnight overflows, by between 1.5 and 1.8 times. The three days that fit are two weekends and the Monday of the Dragon Boat holiday. Averaged across the fortnight the registry published 552 announcements a day; averaged across working days only, closer to 810.

The practical shape of that failure is worth being precise about, because it is not a crash. A day-slicing crawler asks for page 51 of 13 June 2025, receives HTTP 404, and concludes it has reached the end of that day. It has reached the end of what it is allowed to see: 500 of 889 records. It records a successful day. It moves on. The 389 announcements it did not collect leave no trace anywhere in its logs.

The year it started mattering

So we asked when a day stopped fitting. Sampling the same week of June in each year of the registry's life:

Sampled dayAnnouncements that dayFits in the 500-record window?
13 June 20173Yes
12 June 2018~40Yes
11 June 2019~60Yes
9 June 2020~140Yes
8 June 2021~190Yes
14 June 2022~360Yes — with little room left
13 June 2023~610No
11 June 2024~620No
10 June 2025821No
9 June 2026~690No
11 August 2026~780No

These are single sampled days rather than annual averages, and the figures marked with a tilde are derived from page counts rather than read from a count header, so treat the curve as a shape and not as a series. The shape is unambiguous anyway: day-slicing was sufficient through 2022 and insufficient from 2023.

That is the finding we would want if we were buying. A vendor's China bankruptcy coverage being complete in 2021 tells you nothing about 2024, and the transition needed no announcement, no version bump and no deprecation notice on either side. The registry simply got busier than the window a crawler was given.

It is still getting busier. Yesterday the registry advertised 90,746 pages. Today it advertises 90,838 — about 920 new announcements in twenty-four hours, against a reachable window that has stayed at 500 the whole time. Whatever the gap is, it widens by roughly two windows a day.

The axis that does get underneath the wall

Yesterday we tested the two structured filters this interface exposes — publication date and case type — found that neither shrank a busy day below the window, and concluded the chain was not enumerable. We flagged one axis as untested: a free-text field matched against the announcement title. That was the axis that mattered, and the conclusion has to be withdrawn.

It works, and it works for a reason specific to Chinese court practice rather than to this website. Announcement titles on this registry almost all begin with the court case number, in the standard form used across the PRC court system: year in brackets, then a court code, then a character indicating bankruptcy, then a sequence number and a closing character. The court code opens with a one-character provincial abbreviation. That is not free text in any meaningful sense — it is a closed alphabet of about thirty-two values embedded in a field the interface will let you search.

So the free-text field turns into a province facet. On 13 June 2025 — the busiest day in the fortnight, 889 announcements across 89 pages, of which 50 pages are reachable — the province grid looks like this:

MeasureValue
Non-empty province cells that day31 of 33 characters tried
Largest cell232 records / 24 pages
Cells exceeding the 50-page wall0
Sum of cell record counts925
True day total889
Gross duplication36 records (4.0%)
Requests to cover the day this way33 census probes + 109 page fetches = 142
Records reachable without the axis500 of 889 (56%)

Every cell fits, with the largest at less than half the wall. The 4 per cent duplication is a provincial character appearing somewhere else in a title — typically inside a company name — and it is cheap: you pay for it in a few extra fetches and remove it on the identifier.

So the answer to the question we left open yesterday is yes. The collection is enumerable. What follows is the part we did not expect.

The axis leaks, and you cannot see it leak

Having an axis is not the same as having a partition. To find out which one this is, we needed a day where we already knew the complete answer, which meant a day small enough to walk without any filter at all. 8 June 2025, all 46 records, every identifier known.

Running the same thirty-two-character province alphabet against that day and walking every non-empty cell recovered 45 of the 46. Four records were returned by two cells each. One record was returned by none.

The record that escaped is a debt-declaration notice from a court in Inner Mongolia. Its case number uses the character . Our alphabet had , which is the abbreviation an outsider would reach for and is not the one the court system uses in case numbers. One wrong character in a hand-built alphabet, one record silently outside every cell we asked for.

Then we checked whether that mistake is detectable from inside, and this is the finding we would most want a buyer to take away. Across the whole of June 2025:

Character queriedRecords returnedWhat it actually is
~80The Inner Mongolia court code
9Not a court code — matches company names
~240The Yunnan court code
10Not a court code — matches company names
(Guizhou)~200The Guizhou court code
~40Not a court code — matches company names

A wrong key does not return an error. It does not return zero. It returns a small pile of genuine, correctly-formatted, verifiable announcements, because those characters really do occur in Chinese company names. Build your alphabet from the abbreviations a reasonable person would guess — and note there is no consistent rule to guess by, since Yunnan uses the character an outsider expects while Guizhou does not — and every cell in your grid returns data. All of it real. None of it complete. Nothing in the responses distinguishes “this province was quiet” from “you asked for the wrong province.”

The same ambiguity sits one level down. Paging past the end of a real slice returns HTTP 200 with a body of about 3,418 bytes and no rows. Querying a string that matches nothing at all returns HTTP 200 with a body of 3,418 bytes and no rows. Those two responses were byte-identical in our runs. “You have finished” and “you asked for something that does not exist” are the same event on the wire.

And the leak is invisible precisely where it matters. We caught it only because 8 June was small enough to establish ground truth without the axis. On 13 June — a day you actually need the axis for — there is no unfiltered walk to compare against, so a missing province would present as a slightly lower total and nothing else. The one day you can audit is the one day you did not need to.

What is not the bottleneck

One thing we had left open for two days: whether the announcement bodies are reachable once you can name them. They are. Detail pages for both identifier shapes we encountered returned HTTP 200 carrying the full announcement text — creditor declaration windows, meeting dates, the operative paragraphs — along with links to the original attachments the court filed.

So the constraint on this registry is entirely in the addressing layer. Every announcement you can name, you can read. The difficulty is naming them, and the naming problem has a solution that is invisible from the outside and silently wrong if you get one character of it wrong.

What it costs, honestly

We should say plainly that this is not expensive, because the opposite would be more convenient for us to claim.

Working from the registry's own numbers — 908,378 records across 90,838 pages, ten rows per page, with the axis needed only on days above the window — a full sweep looks like roughly 91,000 page fetches for the data itself, about 3,700 one-per-day probes to find out which days need the grid, about 32,000 province census probes on the post-2022 working days that do, and a few per cent on top for duplicates. Call it 130,000 requests, give or take. At the deliberately slow serial rate we used, that is about five days of continuous, polite crawling on a single connection.

That figure is an estimate built on stated assumptions, not a measurement: we have not run a full sweep and we are extrapolating a per-day grid cost from one day. Treat it as an order of magnitude.

Five days of crawling is not the cost. The cost is the four things this article had to establish before that crawl would be correct: that the wall exists at all; that it stopped binding on your data somewhere in 2023 without telling you; that the title field carries a structured key that gets you underneath it; and that the alphabet of that key is and not , which no response will ever tell you. Then you rerun it, because the registry added about 920 records while you were reading this page.

That is also, in miniature, why coverage claims about Chinese court data are difficult to evaluate from outside. Nobody in this story is behaving badly. The registry publishes everything and serves what it serves. The crawler asks correctly and receives HTTP 200. Both sides are working exactly as designed, and the result is a corpus that is quietly missing a third of its working days from 2023 onward with no error anywhere in the chain.

What we did not test

And the standing caveat on our own side: our corpus has its own documented gaps — thin coverage in the most recent years, a large number of empty case-type fields, court-name variants that are not normalised. This page argues something commercially convenient for a company that sells Chinese legal data, which is exactly why every number in it is one you can reproduce yourself in an afternoon.

Frequently asked

How complete is China's national bankruptcy announcement registry through its own interface?

As of 2 September 2026 the registry states that it holds 908,378 announcements. Its list endpoint stops serving pages after page 50 at ten rows per page, so an unfiltered crawl reaches 500 records, or 0.055 per cent. Filtering by publication date raises that substantially, and adding a provincial axis drawn from the case number in each announcement title raised it further — on the busiest day we measured, from 500 of 889 records to a grid whose every cell fitted under the wall.

When did date-slicing stop being enough to enumerate this registry?

Between 2022 and 2023, on our sampling. A single sampled day in June 2022 carried roughly 360 announcements, inside the 500-record window. The equivalent day in June 2023 carried roughly 610, outside it. Nothing about the interface changed; the publication rate grew past the window. A pipeline built before 2023 that paged each day to exhaustion was complete when it shipped and began dropping records silently afterwards, receiving HTTP 200 throughout.

What does a pagination cap look like in an ETL pipeline's logs?

Like a successful run. Asking for page 51 of a day with 889 announcements returns HTTP 404 with a 1,148-byte body, which a day-slicing crawler reads as the end of that day. It records the day as complete with 500 of 889 records. Separately, paging past the genuine end of a small slice returns HTTP 200 with a roughly 3,418-byte body and no rows, which is byte-identical to what you get for a query matching nothing at all — so “finished” and “no such thing” are indistinguishable on the wire.

Can a free-text title filter be used to enumerate a capped Chinese court list?

Yes, and not by guessing strings. Announcement titles on this registry begin with the court case number, whose court code opens with a one-character provincial abbreviation drawn from a closed set of about thirty-two values. Querying that character partitions a day into cells that all fitted under the pagination wall on the busiest day we measured, at about 4 per cent duplication because those characters also occur in company names. It is an axis, not a partition.

How would you know if your provincial axis was missing a province?

From inside the interface, you would not. On a day we could enumerate completely, a thirty-two-character alphabet recovered 45 of 46 records; the miss was an Inner Mongolia case number, which uses one character where our alphabet had another. Querying the wrong character does not return an error or an empty result — it returned nine genuine announcements over a month, because that character appears in company names. Every cell comes back with real data and the grid still leaks. The only way we detected it was having ground truth from an unfiltered walk, which is only possible on days small enough not to need the axis.

Bring us the slice you cannot reconcile.

SinoVerdict licenses PRC judgment data to teams building legal AI — bulk delivery, a REST API and an MCP endpoint over a 160M+ record corpus, English-indexed, with case-number citations to the original judgments and the known gaps written down before anything is signed. If you have a date range where your counts and someone else's do not agree, send it over and we will run the comparison with you. Write to chenjiaxin@wenshucha.com or use the form.

Request trial access