// Product
From diff to decision, with a record of why.
robustiq compares a base and a head ref, reads your CI test history and scores the change for risk, confidence and evidence coverage. Then it checks the change against your policy gates. You get one decision that names the failed gates and the risks nothing has verified, and you can record it to a durable log to query later.
- 1Compare two refsbase main .. head feature/payment-retry
- 2Read CI test historycheckout.spec › retries once on 503 · flake score 64
- 3Score the changerisk 62 · confidence 71 · coverage 58%
- 4Name what nothing verified2 top risk drivers · 3 unverified risks
- 5Check your policy gates4 gates · 2 failed: evidence coverage, unverified risks
- 6DecideHOLD
- 7Keep the recorddecision log · evidence bundle · notification
// The full report
The verdict and the numbers behind it, on one page.
Ask for the full format and the summary gets a policy gate table with each gate's actual value, threshold and result. Here is the full report for one pull request.
- 01Retry fires twice when the gateway times out
- 02Idempotency key not reused on the second attempt
- 03Refund path after a partial capture
| Policy gate | actual | result |
|---|---|---|
| risk ≤ 70 | 62 | PASS |
| confidence ≥ 60 | 71 | PASS |
| evidence coverage ≥ 70% | 58% | FAIL |
| unverified risks ≤ 1 | 3 | FAIL |
Both score gates pass. Evidence coverage and unverified risks hold the pull request back, and the report says so.
// Risk and confidence
Two scores, because risky and unknown are different problems.
A change can be risky and well tested, or harmless and barely tested. robustiq keeps the two numbers apart, so you know whether to slow down or to gather more evidence.
How risky the change is.
What the diff touches and how much it changes. The report names the top risk drivers: the files and behaviours pushing the number up.
How much evidence you have about it.
How much of what changed is backed by tests that exercise it and a CI history you can trust. Low confidence means nobody has checked the code yet, not that it's bad.
#482 sits in the ship corner on scores alone but still gets a HOLD. Evidence coverage and unverified risks failed their gates. The scores feed into the decision, but your gates make the call.
// Unverified risks
What nothing has checked, named before the merge.
Evidence coverage tells you how much of a change is covered. Unverified risks tell you what isn't, one behaviour at a time, so a reviewer knows what to ask and a coding agent knows which test to write.
- Listed, not just counted.Each item names one behaviour the diff changes that nothing has checked.
- Gate them if you want to.A gate such as unverified risks ≤ 1 holds the change until the list is short enough. PR #482 has three.
- Readable by agents.
robustiq_reportreturns the same list over MCP, so an agent can close the gaps before it asks for review.
- Retry fires twice when the gateway times outpayments/retry.ts › captureWithRetry()Suggested check
Time out the gateway after the first attempt; assert exactly one capture. - Idempotency key not reused on the second attemptpayments/retry.ts › buildRetryRequest()Suggested check
Assert the second attempt sends the first attempt's idempotency key. - Refund path after a partial capturepayments/refunds.ts › refundCapture()Suggested check
Refund an order whose capture was partial and retried; assert the refunded amount.
// Policy gates
You set the thresholds. robustiq measures.
A gate pairs one measured value with a threshold your team chooses. The full report puts each gate's actual value next to its threshold, so a held pull request comes with the number that held it.
| Gate | Rule | Threshold | ||
|---|---|---|---|---|
| risk | ≤ | 70 | ||
| confidence | ≥ | 60 | ||
| evidence coverage | ≥ | 70% | ||
| unverified risks | ≤ | 1 | ||
| add gate | ||||
- 01
Actual next to threshold.
The report prints each gate's actual value and threshold, plus PASS or FAIL.
- 02
Failures are named.
A HOLD or NO-GO always names the gate that failed. Nobody has to reverse-engineer a score.
- 03
One policy for CI and agents.
An agent asking over MCP is checked against the same gates as your pipeline.
// Decision log & Trends
Every decision on record, and which way they're trending.
Recorded decisions go into a durable log you can filter by branch or pull request. The dashboard's Trends tab charts risk and confidence across that log, so you can spot a slow slide before it turns into an incident.
| PR | Decision | Risk | Conf. |
|---|---|---|---|
| #482 | HOLD | 62 | 71 |
| #481 | GO | 24 | 82 |
| #480 | NO-GO | 81 | 38 |
| #479 | GO | 18 | 88 |
| #478 | HOLD | 47 | 56 |
robustiq_decision_trend by branch or pull request.// Evidence bundles
Audit-ready: the decision and its evidence, in one bundle.
For a recorded decision, robustiq can write an evidence bundle: the verdict, the gate results, the risk drivers and the unverified risks it was based on. You attach it to a change approval, and open it again in a post-incident review.
Change approval
Attach the bundle to the change request. The approver sees the evidence the gates saw, not a screenshot of a green build.
Post-incident review
Open the bundle for the release in question: what was known at merge time, which gates passed, and what nothing had verified.
- evidence/pr-482/
- ├─ decision.jsonHOLD
- ├─ gates.json4 gates · 2 failed
- ├─ risk-drivers.json2 drivers
- ├─ unverified.json3 items
- └─ inputs.sha256
{
"pull_request": 482,
"base": "main",
"head": "feature/payment-retry",
"decision": "HOLD",
"risk": 62,
"confidence": 71,
"evidence_coverage": 0.58,
"failed_gates": ["evidence_coverage", "unverified_risks"],
"unverified_risks": 3
}// Notifications
The right people hear about a HOLD.
robustiq can send a notification when a decision is recorded, so a held pull request doesn't sit unnoticed in a CI log.
// Deployment & data
What robustiq reads and what it keeps.
Each decision is computed from two sources (your git repository and your CI test history) and checked against the policy you set. Here's what that means for your code and your data.
Your repository and your CI history.
- Git repository
- to compare a base ref with a head ref
- CI test history
- ingested results for evidence and flake scores
- Your policy
- the gates and thresholds your team sets
What it keeps.
- Decision log and evidence bundles
- for the decisions you record
- MCP calls
- read-only: no records, bundles or notifications