"It still works" is an opinion. Tallystone turns it into evidence. Here is exactly how we prove two versions of a system behave the same — and surface the places they don't.
You export data and screens twice: once from your current version (the baseline), once from the upgraded target version. That's database extracts — one file per table — and screenshots of the screens that matter. You deliver them through a private, encrypted channel. Tallystone never needs a connection into your live environment, and never sees more than the two snapshots you choose to send.
Before any AI is involved, we compute a precise structural diff. This is ordinary, repeatable computation — the same inputs always produce the same result:
The diff is exhaustive and exact. Nothing is sampled, nothing is guessed. If a single premium value moved, we have it.
An exact diff of a real system produces thousands of differences, and most of them are supposed to be there: new version metadata, re-sequenced identifiers, timestamps, configuration for new features. Handing all of that to a review team buries the real problems.
So Tallystone has a frontier AI model read every difference and classify it: an expected upgrade change, or a likely regression — with a written rationale for each call. Screens are compared by a vision model that ignores cosmetic reflow and version stamps and looks for functional change: a missing field, an altered value, a button that vanished.
The AI proposes; it never decides. Its job is to sort the signal from the noise so your people spend their time only where judgement is actually needed.
Every difference becomes a finding in your portal, ranked by severity. Your team works each one:
Every status change and comment is recorded against the finding. When the project closes, you don't just have a working upgrade — you have a defensible record showing exactly what changed, what was reviewed, and who signed it off.
The diff is deterministic, so it's trustworthy and auditable. The AI only triages what the diff already found — it can't invent a difference or miss one, because it never does the finding. And the final call always belongs to your team. Exactness where it counts, judgement where it counts, and a person accountable at the end.