You say what a visitor wants
One intention per scenario, ending with "Stop when…", plus what the page must show.
Say what a visitor wants in plain words. Jev, a small, fast model, clicks through a real Chrome to do it, on desktop and on a phone. QAJev judges the result from the page itself, never from the model's opinion.
$ uv tool install git+https://github.com/joseairosa/qajev
Example: QAJev checks that a visitor can find the Pro plan price on a desktop browser and on a phone. Jev clicks Pricing on desktop, and on the phone opens the menu and then clicks Pricing. Both runs show "$29 per month" and pass; the gate is PASS.
No selectors, no recorded clicks to fix after every redesign. A test says what a visitor wants and how you'll know they got it.
One intention per scenario, ending with "Stop when…", plus what the page must show.
Jev reads the page as text and picks one action at a time, in about a quarter of a second each. It types, scrolls and follows links like a person.
Jev saying "done" is only a hint. Your expectations are checked against the real page, then written up in a report you can read.
Every scenario ends with one outcome, and each one tells you whose problem it is. These reasons come from real reports.
Nothing to do.
all 2 check(s) passed
A check failed or the page didn't load. Here the test expected $25 and the page said $29: Jev's "done" didn't fool anyone.
Jev reported DONE but: page shows '$25 per month' (Pricing Starter $9 per month Pro $29 per month …)
Jev found no way forward. That's often a real usability problem worth a look.
the expected text is in the page but 1 screen(s) of scrolling did not bring it on screen
A time or cost budget ran out, or the page never stopped changing. It says nothing about your product.
3 of Jev's moves went stale before they ran (the page kept changing)
Each run ends with one gate, PASS, FAIL or INCOMPLETE, and an exit code your CI understands. Extra problems seen along the way (script errors, failed requests, broken images) are listed separately as findings, from S1 (worst) to S3.
Each website scenario also runs in a phone view: a phone-sized screen, touch and an iPhone user agent, like your browser's device mode. A failure that only shows on the phone is a real mobile bug, and the report says so.
Some bugs only show on the device itself. Add --real-devices ios,android and the same scenarios also run in Safari on the iOS Simulator and Chrome on an Android emulator, on devices QAJev starts for the run and leaves untouched: a throwaway iPhone clone, a read-only Android emulator.
Every run writes a self-contained report.html (light and dark), plus Markdown and JSON for programs. Projects remember the previous run and show what changed.





QAJev is an MCP server too. Your coding agent can crawl a site, check a change the way a user would, read the report and look at the screenshots, then tell you what it found.
# one line, then give it the agent prompt
claude mcp add -s user qajev -- qajev mcp# ~/.codex/config.toml [mcp_servers.qajev] command = "qajev" args = ["mcp"]
// the client's MCP settings { "mcpServers": { "qajev": { "command": "qajev", "args": ["mcp"] } } }
Paste AGENT_PROMPT.md into your agent's instructions. It knows which tool to use, how to write a good check and how to report the result honestly.
Long runs return a job id at once. Any agent, or you, can see progress, what Jev is doing right now and the money spent, and stop it cleanly.
Many agents can share a laptop. Browser runs take turns and say what they're waiting for, so nobody starts a second copy by accident.
Agent runs use QAJev's own invisible Chrome with a fresh profile. Nothing pops up on your screen and nothing is left behind.
qajev top shows what's queued and running on the machine, what Jev is doing at this moment, and what each run costs, down to a hundredth of a cent.
The same plain-language steps work in a native app on a simulator, or in a game's menus while its own bot plays in real time.
Native apps and mobile websites on the iOS Simulator and an Android emulator. QAJev boots a throwaway iPhone clone (deleted afterwards) or a read-only Android emulator, and installs your build if you give it one. The microphone stays deaf.
It also catches what screen readers miss, like tab bars that VoiceOver doesn't announce as buttons.
Godot and Electron games through a small bridge. Jev works the menus and picks the upgrades; the game's own pilot plays in real time; QAJev watches frame rate, memory, soft-locks and errors.
Form posts and deletes are blocked inside the page. Tests that change data only run on your own machine.
Sign out, delete, pay, billing, "close all"… Jev never sees them, so it can't press them.
Secret and payment fields are disabled. For signed-in areas, you sign in once yourself in QAJev's own browser.
Every run has one ($1 unless you say otherwise), checked before every paid call.
Pages that listen hear only what your test says. On simulators, microphone prompts are refused.
Every download is refused. The smoke crawl checks download links without fetching the file.
QAJev runs its own Chrome with its own profiles. Your everyday browser and accounts are never touched.
The smoke crawl reads robots.txt and waits between pages. QAJev is not a load tester.
You need a Mac or Linux machine, Chrome, and one key for Jev from TypeSafe or OpenRouter. The smoke crawl doesn't even need the key.
uv tool install git+https://github.com/joseairosa/qajevmkdir -p ~/.qajev && echo 'OPENROUTER_API_KEY=sk-or-...' >> ~/.qajev/.envqajev doctorqajev smoke http://localhost:3000qajev check http://localhost:3000 --goal "Open the pricing page. Stop when the prices are visible." --expect-url /pricing