Kyber Cypher

Field Logs

Field Log 033

Four of thirty

Field Log // 033 Status Live Difficulty Free Cost Nothing
The Story

I wrote something to read a shop's catalogue and build a product grid from it. It ran, it produced clean output, nothing errored, and it found four products. The shop has thirty.

The code was not broken. It was pointed at the wrong page, and the wrong page looked exactly like the right one.

The site is the modern kind: the server sends a shell and a small featured selection, and the rest arrives afterwards when the page runs its own code in the browser. Fetching the address gives you what the server chose to include, which was four items. A browser would then have filled in the other twenty six. My fetch never ran that step, because a fetch is not a browser, so it took the four and reported success.

Photographing a shop window is not an inventory. The window is a selection someone made this morning, and the stockroom is not in the picture.

The specific failure is interesting, but the thing worth carrying away is the failure mode: a scrape that returns plausible, well formed, internally consistent data that is simply incomplete. There is no error to catch. The only thing that exposes it is knowing, independently, roughly how many things there should be.

Finding the rest meant finding where the catalogue really lives. The paginated listing turned out to be a dead end, because its page control was built in a way that the server never honours, so every variation of the address returned the same first batch. What did work were the category pages, each of which server rendered its full membership. The union of those, deduplicated, was the catalogue.

Then I did the thing I should always do and had not: I counted it a second, independent way, and then a third. The listing's own pagination control implied a total. The site's machine readable index implied another. When all three agreed, I believed the number. They did agree, and they disagreed with my first answer by twenty three products.

The Build

How to find a page's real data source, combine partial sources into a complete set, remove duplicates safely, and verify a count independently before you trust it. Read only throughout; this is about reading public pages you are entitled to read, politely.

1. Establish the expected count from outside the scrape

Do this first. Without an independent expectation you have no way to detect a plausible shortfall, which is the whole failure mode.

# acceptable sources for "how many should there be"
#   an admin screen you own
#   a count the site states itself
#   the site's machine readable index
# write the number down BEFORE you look at your output

2. Compare what the server sends with what the browser shows

One command tells you whether content is server rendered or filled in later.

# what the server actually sent
curl -s '<listing-url>' > server.html
grep -c '<marker-for-one-product>' server.html

# then count what you see in a real browser on the same page
# a big gap means the rest arrives client side

If the numbers differ, stop treating the page as your data source and go looking for where the data comes from.

3. Find the real source in the network panel, not by guessing

Open developer tools, reload, and watch what the page requests. The data is usually one of three things.

# look for, in order of usefulness
#   1. a request returning JSON with all the items  -> use that
#   2. an embedded blob in the HTML (a script tag of state)
#   3. per-category or per-section pages that ARE server rendered

# and check the obvious indexes while you are there
curl -s '<site>/sitemap.xml' | grep -c '<loc>'

4. Test pagination properly before trusting it

A page control drawn in the markup does not mean the server accepts a page parameter. Try the variations and compare results.

# do these return DIFFERENT items, or the same batch each time?
curl -s '<listing>?page=2'   | grep -c '<product-marker>'
curl -s '<listing>?offset=24' | grep -c '<product-marker>'

# compare the actual identifiers, not the counts. Same count with the
# same ids means the parameter is being ignored.

Mine were ignored, which is why the first attempt plateaued at one batch and looked finished.

5. Union the partial sources, deduplicating on a real identifier

Combine everything you found. Deduplicate on the item's own identifier rather than on its title, because titles repeat and identifiers do not.

found = {}
for source in [listing] + category_pages:
    for item in parse(fetch(source)):
        found[item.id] = item          # id as the key, so dupes collapse

# and log the walk, so a future run can show HOW it got there
#   /listing            24 items,  24 new,  running 24
#   /category/one        1 item,    0 new,  running 24
#   /category/two       21 items,   2 new,  running 26

That running total is worth keeping. It shows which source contributed what, which is how you notice when a source stops working.

6. Verify the count two independent ways, then a third

Agreement between methods that could fail differently is what justifies belief.

# three methods that do not share a failure mode
#   A. union of listing + categories, deduplicated
#   B. what the pagination control implies (batches x per-batch)
#   C. the machine readable index, plus anything it is missing
# all three agree -> trust it. any disagreement -> investigate, do not average.

7. Put a shrink guard on it, and make it refuse

Once you know the correct count, defend it. A future run that finds fewer items is far more likely to be a broken scrape than a real removal.

# before writing new output over good output
if new_count < stored_count and not override_requested:
    refuse()        # keep the good data, exit non-zero, say both numbers

# and test the guard by faking a shortfall, so you have SEEN it refuse

Then be honest about what remains. Mine settles at a number lower than the shop's own total, because some products are not published and a public page genuinely cannot show them. A gap you have explained is fine. A gap you have not looked at is the bug.

The honest catch: category pages and embedded state are implementation details of someone else's site, and they change without warning. Expect this to break, make it break loudly, and check the count rather than the exit code.

Related: the API price is not the price is the same distrust pointed at a field instead of a page.