Kyber Cypher

Learn

Field Log 032

The API price is not the price

Field Log // 032 Status Live Difficulty Free Cost Nothing
The Story

I was building a product grid for a shop. The platform holding the products has a proper API, with a field called price. I took the price from the price field, which is the single most reasonable thing a person can do, and it was wrong on a third of the catalogue.

Not slightly wrong. Wrong in the direction that matters: lower. One item would have been listed at thirteen seventy five while the checkout charged eighteen sixty. Nine of twenty seven products were affected.

The field was real and the number in it was real. It was the pre markup cost, the figure before the storefront applies its own pricing. There was a cost field too, sitting next to it, which is what made me look twice: if price is what the customer pays, why does the object also carry a cost, and why is the gap between them not the gap I expected? I pulled up the public page for that product and the public page disagreed with the API.

A wholesale invoice and a shelf label both state a price for the same tin. Both are correct. Only one of them is what the person at the till pays.

What saves you here is not scepticism about APIs in general. It is a narrower habit: when money is involved, take the number from the surface the customer actually sees, and treat anything else as a hint. The storefront page is not a more authoritative data source in any technical sense. It is simply the thing that is true by definition, because it is what the buyer reads and what the checkout honours.

The deeper version of the lesson is about what "authoritative" means. An API is authoritative about its own model. It is not automatically authoritative about your question. The field was not lying. I had asked it something it was never answering.

I kept both numbers in the end. The customer facing price is what gets published, and the API number sits beside it in the data as evidence, so the next person who wonders why they differ can see that somebody already checked.

The Build

How to cross check a catalogue against the live storefront, and the general pattern for validating data that looks authoritative. Applies to any integration where the number has consequences.

1. Write down which surface is the truth, before you write any code

One sentence per field. This takes two minutes and it is the design decision that prevents the whole class of bug.

# sources.md
#   price     -> the public product page. the customer pays this.
#   title     -> API. it is the canonical name.
#   link      -> API, but only the field that is a real URL
#   image     -> API, highest resolution available
#   stock     -> ? decide, and say why

Where a field exists in two places, name the winner explicitly rather than taking whichever is convenient.

2. Read the whole object once, with your eyes

Before mapping anything, print one complete record and look at every key. Neighbouring fields are the clue that you are about to pick the wrong one.

# dump one record, formatted, and read it
curl -s -H "Authorization: Bearer $TOKEN" <endpoint> | python3 -m json.tool | head -60

# the question to ask on every money field:
#   is there another field nearby that could ALSO be called a price?
#   if yes, you do not yet know which one you want

3. Compare against the customer facing page for a sample, by hand

Pick three products, open their public pages, and compare. This is five minutes and it is the step that caught mine.

# what the API says
#   product A: price field = ............
# what the public page shows
curl -s '<public-product-url>' | grep -oE '<whatever marks the price on that page>'

# if these differ on even one product, stop and find out why
# before building anything on top of the API number

4. Take the real number from the page you are already fetching

If you verify that a product link works, you are fetching the page anyway. Read the price out of that same response and the correct number costs you nothing extra.

# one request, two results
status, body = fetch(product_url)
if status != 200: exclude_this_product()
price = extract_price(body) or fall_back_to_api_price()

# and record WHICH source each price came from
#   price_source: "storefront" | "api-fallback"

Recording the source is what lets you audit later without re-deriving anything. If a price came from the fallback, you know to treat it as unverified.

5. Keep the disagreement in the data, not in your memory

Store both numbers. A future reader who finds them different has the evidence that it was deliberate, rather than discovering a mystery.

# per product
#   price        : what we publish  (from the storefront)
#   api_price    : what the API said
#   price_source : which surface won

6. Make a mismatch a build failure, not a warning

Once you know the two sources can disagree, check it on every run. Money is the category where a warning nobody reads is the same as no check at all.

# in the build or the nightly refresh
#   for each product: published price == storefront price ?
#   any mismatch -> non-zero exit, do not deploy, say which products

7. Verify every outbound link before publishing it

While you are there. A catalogue page whose buy links fail is worse than one product missing, and it is the same loop.

# every link you are about to publish, not a sample
#   request it, require 200, exclude anything else
#   and never construct a URL from an identifier unless you have
#   confirmed that identifier belongs to the public-facing namespace

That last clause is its own trap: internal identifiers and public ones often look similar and are not the same, so a URL you assembled yourself can be well formed and still lead nowhere.

The honest catch: this makes your page agree with the storefront at build time. Prices change on the platform without telling you, so the real guarantee comes from how often you re-run the check, not from the check itself.

Related: the scrape that found four of thirty is the same distrust applied to how much of a catalogue you are actually seeing.