Intent-driven, automated, visual quality assurance has arrived.
Were the site changes you requested what was deployed? The only AI that you want at this stage is one that cannot make things up.
Twenty sites, thirty, forty. Every one of them is a business somebody depends on, and every one of them has your name on it.
A client asks for a new phone number, a new photo, a new heading. An AI agent makes the edit. It is fast, and it is cheap, and it works most of the time.
A person used to look. Now nobody does, because there are too many pages and too many changes. The agent does not look either. It cannot.
A phone number that stopped working. A photo of the wrong person. A button that quietly disappeared. You did not make the change, but you take the call.
A chat AI will describe the page and sound confident. It can also make things up, miss what matters, and tell you it looks fine. Now you have two things that can be wrong.
We build and run client sites too. One rebuilt site passed every check we had and still had 128 things a visitor would see. We took that call.
Not another AI that looks and describes. A machine that compares, measures and asks only yes-or-no questions. It is called the GntLit, and it runs on your own computer.
Every pixel. Every word. Every link, phone number, picture and form. Code does this, not an AI, and code has no opinions. Anything that changed is marked. We say it is lit.
Is the phone number still there? Is a button missing? Is this the change that was asked for? The AI answers each question with a confidence score. It cannot write a sentence, so it cannot invent anything.
We planted 791 known mistakes in our own site and checked whether the AI caught them. It catches a missing phone number every time. It misses a small typo more than half the time. We publish both numbers.
The change passes, or a person needs to look, or the change is blocked. Beside it, a dated folder with the pictures, every difference in words, and every answer with its score. You can show it to the client.
A person moves the change, or your own publishing steps do. Nothing leaves your machine. The AI is never trained on your pages. It has no voice.
The gauntlet, lit. Every page takes every check. Only what changed is lit. Nothing is made up.
Keep scrolling to watch a real page run it.
Two lines of inspectors. On the left, code: it checks pixels, words, links, pictures, forms and the hidden code behind the page. On the right, the AI judge, asked only about the spots the code marked. No page skips a station. Keep scrolling.
Animations stopped, every picture loaded, fonts ready. Then the page is captured twice, in two fresh browsers, on your own machine. If the two screenshots disagree anywhere, a slideshow or a clock for example, that spot is painted out and written down. If they agree, you have a trusted copy to compare against.
desktop and phone sizes · both screenshots identicalThe header, the hero, every block and the footer are paired across the two versions by their role, their heading and their first words. If a section has no partner, that is a fact of its own: a section is missing, or a section is new.
7 sections · 7 matched · 0 missingEach matched section is cut into slices one screen tall, with a small overlap so nothing falls between two cuts. The slices are never shrunk. An AI asked to read a shrunken page is reading smudges, and smudges are where invention starts.
this page: 4 slices, all at full sizePixels, exactly. Words, section by section. Links. Phone and email links. Pictures. Forms. The page title and settings. The whole hidden code behind the page. Where the two versions are identical, the blow glances off as frost. Where they differ, the blow lands, and that spot is lit.
found a change: pixels, words, phone links, hidden code. Identical: links, pictures, forms, title.On this page, one slice changed: someone hid the “Call or text Mark” button. The lit box is exactly the button's own box. The same change is written down in words: the text that disappeared, the phone link that was hidden, and the exact place in the hidden code where the two pages part.
the box: 163 by 49 pixels, less than a tenth of a percent of the screenThe judge is Clef, a free AI from Cloudflare that you download and keep. It runs on your own machine. It looks at the lit slice and at a fixed list of questions, such as “are the buttons the same in both pictures?”, and answers each with a confidence from 0 to 1. That is all it can do. We ask it twice, with the two pictures swapped, and keep the less flattering answer.
here: “a button is missing in the new version”, confidence 0.93 one way round, 0.95 the otherBefore any answer counted, the judge sat a test: 791 planted mistakes, each with the right answer written down first. We also checked whether its confidence is honest, on 10,302 answers. When it says 45% it is right about 29% of the time; when it says 76% it is right 87% of the time. So the line where an answer counts comes from that check, not from the number the judge states.
it names a hidden link 100% of the time, a typo 43%, a re-cropped picture 8%. The code found all of them.The result is one of three: the change passes, a person looks, or the change is blocked. If nobody asked for a change, anything lit is unexpected and a person looks. A question is allowed to block a change only after it has passed its test. Every other answer can only hold the change for a person, and says so.
this run: a person looks. The “what changed” questions have passed their test; the “was it asked for” question has not yet.Every run leaves a dated folder on your machine: the screenshots, the painted-out spots, the lit slices, every difference in words, every answer with its confidence both ways round, and the result with the test it rests on. Nothing in it is ever overwritten. Only what changed is lit. Nothing was invented, because nothing here can invent.
the folder: the run, the screenshots, every difference, every answer, and the report pageWe hid the “Call or text Mark” button on our own contact page and ran the whole site. These are the pictures the GntLit saved, cut to the band around the change. The other 32 pages came back unchanged.



We planted known mistakes in our own site and wrote down the right answer before the judge looked. Here is how it did. The misses are printed as large as the catches on purpose. Where the judge is weak, the code still finds the change.
| The planted mistake | Code found it | The judge named it |
|---|---|---|
| A phone or email link hidden (104 planted) | 104 of 104 | 100% |
| A menu item hidden (120 planted) | 120 of 120 | 100% |
| A picture swapped for another (67 planted) | 67 of 67 | 97% |
| One digit changed in a number (60 planted) | 60 of 60 | 78% |
| A one-word typo in a heading (60 planted) | 60 of 60 | 43% |
| A picture re-cropped by 8% (67 planted) | 67 of 67 | 8% |
| Nothing changed at all (60 identical pairs) | 0 false alarms | 3% false alarms |
We also restarted the judge and asked 74 pairs again. Every answer was identical. Swapping the two pictures changed 8% of answers, so every pair is always asked both ways and the less flattering answer counts.
The live site is the reference. The preview is the new version. The note that asked for the change is the yardstick. The report from the latest run will sit here, misses included. There is no report yet, because this page has not shipped a change. The first change will be the first run.
Anything lit that the request does not name goes to a person.
No old version needed. The page is cut into screens and read with fixed questions: a cut-off face, text running into other text, placeholder words left in, no way to call.
The old site is the source. The rebuild is the new version. Count the sections, name every difference, close them one at a time.