GntLit

Intent-driven, automated, visual quality assurance has arrived.

Were the site changes you requested what was deployed? The only AI that you want at this stage is one that cannot make things up.

Scroll
the story · 1 of 13
1 · you

You run websites for other people.

Twenty sites, thirty, forty. Every one of them is a business somebody depends on, and every one of them has your name on it.

2 · what changed

These days, the changes are made by AI.

A client asks for a new phone number, a new photo, a new heading. An AI agent makes the edit. It is fast, and it is cheap, and it works most of the time.

3 · the problem

Nobody looks at the page before it goes live.

A person used to look. Now nobody does, because there are too many pages and too many changes. The agent does not look either. It cannot.

4 · the call

When something is wrong, the client calls you.

A phone number that stopped working. A photo of the wrong person. A button that quietly disappeared. You did not make the change, but you take the call.

5 · the trap

Asking another AI to check does not help.

A chat AI will describe the page and sound confident. It can also make things up, miss what matters, and tell you it looks fine. Now you have two things that can be wrong.

6 · we have been there

We shipped a site with 128 mistakes. Every automatic check said it was fine.

We build and run client sites too. One rebuilt site passed every check we had and still had 128 things a visitor would see. We took that call.

7 · so we built this

So we built a checker that cannot make things up.

Not another AI that looks and describes. A machine that compares, measures and asks only yes-or-no questions. It is called the GntLit, and it runs on your own computer.

8 · how it works, step one

First, code compares the old page with the new one, exactly.

Every pixel. Every word. Every link, phone number, picture and form. Code does this, not an AI, and code has no opinions. Anything that changed is marked. We say it is lit.

9 · step two

Then an AI is asked fixed questions, only about the spots that changed.

Is the phone number still there? Is a button missing? Is this the change that was asked for? The AI answers each question with a confidence score. It cannot write a sentence, so it cannot invent anything.

10 · step three

Every question was tested before it was allowed to count.

We planted 791 known mistakes in our own site and checked whether the AI caught them. It catches a missing phone number every time. It misses a small typo more than half the time. We publish both numbers.

11 · what you get

You get a simple result and a folder of proof.

The change passes, or a person needs to look, or the change is blocked. Beside it, a dated folder with the pictures, every difference in words, and every answer with its score. You can show it to the client.

12 · what it never does

It never puts a change live, never sends a message, and never invents.

A person moves the change, or your own publishing steps do. Nothing leaves your machine. The AI is never trained on your pages. It has no voice.

13 · the name

GntLit

The gauntlet, lit. Every page takes every check. Only what changed is lit. Nothing is made up.

Keep scrolling to watch a real page run it.

the fall · station 00
the fall · station 00

Now you are the page. Watch the checks land.

Two lines of inspectors. On the left, code: it checks pixels, words, links, pictures, forms and the hidden code behind the page. On the right, the AI judge, asked only about the spots the code marked. No page skips a station. Keep scrolling.

station 01 · hold still

First, the page is held still.

Animations stopped, every picture loaded, fonts ready. Then the page is captured twice, in two fresh browsers, on your own machine. If the two screenshots disagree anywhere, a slideshow or a clock for example, that spot is painted out and written down. If they agree, you have a trusted copy to compare against.

desktop and phone sizes · both screenshots identical
station 02 · match

Then the sections are matched up.

The header, the hero, every block and the footer are paired across the two versions by their role, their heading and their first words. If a section has no partner, that is a fact of its own: a section is missing, or a section is new.

7 sections · 7 matched · 0 missing
station 03 · cut

Then each page is cut into screens, at full size.

Each matched section is cut into slices one screen tall, with a small overlap so nothing falls between two cuts. The slices are never shrunk. An AI asked to read a shrunken page is reading smudges, and smudges are where invention starts.

this page: 4 slices, all at full size
station 04 · the code strikes

Eight checks, in order. Four of them find a change.

Pixels, exactly. Words, section by section. Links. Phone and email links. Pictures. Forms. The page title and settings. The whole hidden code behind the page. Where the two versions are identical, the blow glances off as frost. Where they differ, the blow lands, and that spot is lit.

found a change: pixels, words, phone links, hidden code. Identical: links, pictures, forms, title.
station 05 · lit

The diff is on fire. Nothing else is allowed to burn.

On this page, one slice changed: someone hid the “Call or text Mark” button. The lit box is exactly the button's own box. The same change is written down in words: the text that disappeared, the phone link that was hidden, and the exact place in the hidden code where the two pages part.

the box: 163 by 49 pixels, less than a tenth of a percent of the screen
station 06 · the judge

The judge cannot write. That is why you can trust it.

The judge is Clef, a free AI from Cloudflare that you download and keep. It runs on your own machine. It looks at the lit slice and at a fixed list of questions, such as “are the buttons the same in both pictures?”, and answers each with a confidence from 0 to 1. That is all it can do. We ask it twice, with the two pictures swapped, and keep the less flattering answer.

here: “a button is missing in the new version”, confidence 0.93 one way round, 0.95 the other
station 07 · the test

Trust is checked, not felt.

Before any answer counted, the judge sat a test: 791 planted mistakes, each with the right answer written down first. We also checked whether its confidence is honest, on 10,302 answers. When it says 45% it is right about 29% of the time; when it says 76% it is right 87% of the time. So the line where an answer counts comes from that check, not from the number the judge states.

it names a hidden link 100% of the time, a typo 43%, a re-cropped picture 8%. The code found all of them.
station 08 · the gate

At the gate, the result has to show the test it passed.

The result is one of three: the change passes, a person looks, or the change is blocked. If nobody asked for a change, anything lit is unexpected and a person looks. A question is allowed to block a change only after it has passed its test. Every other answer can only hold the change for a person, and says so.

this run: a person looks. The “what changed” questions have passed their test; the “was it asked for” question has not yet.
station 09 · the record

Out the far end with a folder, not a story.

Every run leaves a dated folder on your machine: the screenshots, the painted-out spots, the lit slices, every difference in words, every answer with its confidence both ways round, and the result with the test it rests on. Nothing in it is ever overwritten. Only what changed is lit. Nothing was invented, because nothing here can invent.

the folder: the run, the screenshots, every difference, every answer, and the report page
the pictures you just watched

This was a real run, on our own site.

We hid the “Call or text Mark” button on our own contact page and ran the whole site. These are the pictures the GntLit saved, cut to the band around the change. The other 32 pages came back unchanged.

The live contact page band, with the Call or text Mark button present
the live page, held stillboth screenshots identical
The new version of the same band, with the button gone
the new versionsame slice, same painted-out spots
The diff: the button's own box lit in red, nothing else
the diff, lit163 by 49 pixels
the test · 791 planted mistakes

What the judge catches, and what it misses.

We planted known mistakes in our own site and wrote down the right answer before the judge looked. Here is how it did. The misses are printed as large as the catches on purpose. Where the judge is weak, the code still finds the change.

The planted mistakeCode found itThe judge named it
A phone or email link hidden (104 planted)104 of 104100%
A menu item hidden (120 planted)120 of 120100%
A picture swapped for another (67 planted)67 of 6797%
One digit changed in a number (60 planted)60 of 6078%
A one-word typo in a heading (60 planted)60 of 6043%
A picture re-cropped by 8% (67 planted)67 of 678%
Nothing changed at all (60 identical pairs)0 false alarms3% false alarms

We also restarted the judge and asked 74 pairs again. Every answer was identical. Swapping the two pictures changed 8% of answers, so every pair is always asked both ways and the less flattering answer counts.

what it never does

Six things the GntLit will never do.

this site runs it

Every change to this page will run the GntLit before it goes live.

The live site is the reference. The preview is the new version. The note that asked for the change is the yardstick. The report from the latest run will sit here, misses included. There is no report yet, because this page has not shipped a change. The first change will be the first run.

a client asks for a change

Did only the change my client asked for ship?

Anything lit that the request does not name goes to a person.

an agent builds a page

Would a visitor wince at this page?

No old version needed. The page is cut into screens and read with fixed questions: a cut-off face, text running into other text, placeholder words left in, no way to call.

a site is rebuilt

Is the rebuild the same site the client already has?

The old site is the source. The rebuild is the new version. Count the sections, name every difference, close them one at a time.

Run it again