Kill flakes and debug failures with full per-test history
A flakiness score for every test, computed from your auto-approved builds. The history of each recurring change. A one-click ignore that never hides a new regression. And when a test fails, the Playwright trace and the failure screenshots are already in Argos, for you or the agent fixing it.
See flakiness before it hurts your CI
Argos scores every test from its auto-approved builds, where a change means the test moved on its own. The score, its trend and every change the test produced live on one test page.
A flaky badge on every changed test
Reviewing a build, you see which changed tests are known to be unstable before you decide. Hover the badge for the numbers, select it to open the test page with everything Argos knows about that test.
Every change, grouped by shape
A change is one exact visual difference, not one occurrence of it. Argos groups every diff that looks the same, so the recurring change stands out from the one-offs, with its occurrences, first seen and last seen, each linking to the build.
Ranked by flakiness
The Tests dashboard sorts a project's tests flakiest first; the account-wide one does the same across every project you can see. The first page is the backlog, for you or an agent.
Learn moreOnly auto-approved builds count
A change on a pull request is usually intentional, so scores use auto-approved builds only, over a period you pick: 24 hours to 90 days, 7 days by default, kept in the URL.
Learn moreFive metrics, one score
Flakiness from 0 to 100, plus Builds, Changes, Stability and Consistency, broken down over time. A cliff in the chart usually points at the commit that introduced the flake.
Learn moreSilence noise without hiding regressions
Argos treats flakiness as technical debt to fix, not noise to ignore. When a change genuinely cannot be made deterministic, ignore it: the ignore covers that one change, and everything else on the screenshot stays under review.
Ignore in one click
From the build page or the test page, select Ignore next to the change, or press I. No config file, no baseline to rewrite. Argos stops asking about that exact change on future builds.
A fingerprint, not a screenshot
Each change carries a fingerprint computed from the shape of its diff. An ignore is a test-plus-fingerprint pair: it absorbs antialiasing noise, but a different change to the same screenshot is still reported. The fingerprint is part of a change's identity in the API and CLI.
Auto-ignore recurring changes
Let Argos ignore a change once it has appeared N times in auto-approved builds over the last 7 days. The default is 3. Auto-ignored changes are attributed to the Argos bot and carry an Auto badge.
Learn moreAn Ignored page you can audit
Every ignored change, with how many builds it has absorbed since and when it was last seen. An ignore that went quiet is a blind spot: unignore it in one click.
Learn moreVisible on the pull request
The PR comment counts ignored screenshots next to the others, as in 4 changed, 3 ignored. The whole feature is a per-project toggle in Project Settings → Flaky detection.
When a test fails, the evidence is already there
The Argos Playwright reporter uploads failure screenshots and traces with your build. You open them in Argos instead of downloading CI artifacts, whether the test failed for good or passed on retry.
Replay the Playwright trace in Argos
Set trace: "on-first-retry" and add the Argos reporter. The trace of a failing test is uploaded with the build and opens in the trace viewer straight from Argos, so you or your agent start where the failure happened.
Every failure, every attempt
With screenshot: "only-on-failure", each failed attempt uploads its screenshot, with retries and retry in its metadata. Failures that passed on retry sit in their own section, so you can inspect a flaky run without it cluttering the review.
One reporter, two lines of config
Add @argos-ci/playwright/reporter and the two use options. The same reporter uploads your screenshots, so debugging and visual testing share one build.
Cypress too
The Cypress SDK gives you visibility on test failures alongside its stabilization: fonts, images and aria-busy loaders settled before capture. Traces are a Playwright feature.
Reproduce a flake on purpose
Run playwright test --repeat-each 5 to run each test several times in one go and catch discrepancies before they reach your main branch.
A thread that outlives the build
Some things belong to a test, not to one run: flaky since the carousel landed, waiting on the upstream fix, safe to ignore until Q3. The test page has its own Activity thread, and a prompt to hand the whole investigation to an agent.
Comment on the test, not the run
The Activity section is the test's own conversation, separate from any build review. Markdown, / commands and @mentions; reply, react and resolve. A comment on a test posts immediately: there is no pending review to batch it into.
Fix with AI
Copy a prompt that names the test, carries the flakiness Argos measured, and tells the agent how to pull the recurring changes and their screenshots before it edits anything. Paste it into Claude Code, Codex, Cursor or any agent with your repository checked out.
Follow a test
The bell in the Activity header notifies you of new comments. Commenting follows the test for you; argos test subscribe does it from the CLI.
From the CLI, API and MCP
argos test comment create posts to the thread; list, resolve and react are there too. The same operations exist over the REST API and as MCP tools, so an agent can leave its findings where the team reads them.
A playbook for humans
Dates and times, fonts, GIFs, background images, loading states: the stabilization guide covers what the SDK cannot infer on its own, for you or the agent you point at it.
Learn moreHand the flake to an agent
Everything on the test page is readable from a terminal. An agent lists the flakiest tests, reads one test's metrics and its recurring changes with their screenshots, finds the non-determinism in your repository and fixes it. What it cannot fix, it ignores.
# Flakiest tests first: the backlog
argos test list --project acme/app --limit 20 --json
# One test: metrics, trend, first and last change
argos test get <testId> --json
# Its changes, most frequent first, with screenshots
argos test changes <testId> --json
# Silence what cannot be fixed (personal access token)
argos change ignore <changeId> --project acme/appCLI
A project token reads tests and changes; ignoring a change needs a personal access token. npx skills add https://argos-ci.com installs the argos-cli skill, which covers the flakiness commands and the token rules.
MCP server
getTest and listTestChanges return the same data at https://mcp.argos-ci.com. Reading needs projects:read; ignoring a change needs reviews:write.
REST API
GET /projects/ and …/changes carry the evidence. Review the agent's fix like any change: a test that stops asserting is worse than a flaky one.
“If Argos fails, it almost always means something visually changed. That signal is incredibly reliable.”
Frequently asked questions
How does Argos know a test is flaky?
Only auto-approved builds count. A change on a pull request branch usually means someone changed the UI on purpose; a change on an auto-approved build means the test moved on its own. Over a period you pick, from 24 hours to 90 days and 7 days by default, Argos measures the test's Stability (the share of builds where it did not change) and Consistency (the share of its changes that were one-offs), and derives a Flakiness score from 0 to 100. Higher is worse.
The score shows as a badge next to every changed test in a build review, and the Tests dashboard ranks a whole project, or your whole account, by it.
Does ignoring a change hide real regressions?
No. Each change carries a fingerprint, a stable signature computed from the shape of its diff, and an ignore is a test-plus-fingerprint pair. The fingerprint absorbs pixel-level noise such as antialiasing but distinguishes genuinely different changes, so a new regression elsewhere in the same screenshot is still reported.
The Ignored page lists every ignored change with how many builds it has absorbed since and when it was last seen. An ignore that went quiet is a blind spot: unignore it in one click and Argos asks for review again the next time the change appears.
Can Argos ignore flaky changes automatically?
Yes, per project. In Project Settings → Flaky detection, turn on Auto-ignore flaky changes and set the minimum number of occurrences. A change counts as flaky once it has appeared that many times in auto-approved builds within the last 7 days; the default is 3. Auto-ignored changes are attributed to the Argos bot, carry an Auto badge on the Ignored page, and can be unignored like any other.
The same settings page has a single toggle that turns the whole ignore feature off for a project: new builds then ignore nothing and the Ignore button disappears from reviews.
How do I see a Playwright trace in Argos?
Add @argos-ci/playwright/reporter to your Playwright config and set trace: "on-first-retry" and screenshot: "only-on-failure" in use. The reporter uploads failure screenshots and traces with the build, and you open the trace viewer from Argos, with no artifact to download. Each uploaded snapshot, traces included, is limited to 50 MB.
Traces are a Playwright feature. The Cypress SDK gives you visibility on test failures alongside its stabilization.
Can an AI agent fix a flaky test?
Yes. The test page has a Fix with AI card with a prompt to copy into Claude Code, Codex, Cursor or any agent working in your repository. It names the test, carries the flakiness Argos measured, and tells the agent to pull the recurring changes and their screenshots with argos test get and argos test changes (or the getTest and listTestChanges MCP tools, or the REST API) before editing anything.
The agent then looks for the usual non-determinism, such as animations, dates, fonts, unordered data and loading states, and fixes the root cause; a change that genuinely cannot be made deterministic gets ignored instead. A project token is enough to read; ignoring needs a personal access token. Review its change like any other: a test that stops asserting is worse than a flaky one. See Fix flaky tests with AI agents.
One platform, from pull request to merge
Stabilize is the last step of the flow. Argos also deploys every pull request to a preview URL, diffs everything it changed, and gives humans and agents one place to review it.
Diff
Visual regression testing for screenshots, and diffs for any file: Markdown, JSON, and more.
Explore DiffSupercharge your product quality
See every change your team and your agents make. Review with confidence, and merge faster.