Can do

Register eval runs as EEE-shaped JSON records, compare any two runs metric-by-metric, and fail CI when a metric regresses.

Best for

Teams shipping AI features who want regressions caught by a gate, not by users.

Screenshots

XingAI Eval Registry
Regression gate diff between two eval runs

Features

  • Every Eval Ever record schema
  • One JSON file per record, git-diffable
  • Metric direction semantics (higher/lower is better)
  • diff --fail-on-regression CI gate
  • Private data directory override
  • MIT licensed

Versions & pricing

Free

$0

  • Every Eval Ever record schema
  • One JSON file per record, git-diffable
  • Metric direction semantics (higher/lower is better)
  • diff --fail-on-regression CI gate
  • Private data directory override
  • MIT licensed
Get started

Enterprise

Contact us

  • Every Eval Ever record schema
  • One JSON file per record, git-diffable
  • Metric direction semantics (higher/lower is better)
  • diff --fail-on-regression CI gate
  • Private data directory override
  • MIT licensed
Contact sales

Want the full source code?

Fork the open-source Next.js app — inventory scan, meal recommendations, and cooking steps. Deploy your own instance on Vercel.

View source on GitHub

Need a custom version?

We build tailored versions for teams and businesses. Tell us what you need and we’ll scope it together.

Contact us

Release roadmap

What’s coming next for this product.

  • ShippedEEE-shaped records + CLI (add/list/show/diff)
  • ShippedFirst consumer: Evidence Engine
  • In progressCI regression gate in consumer repos
  • PlannedEvaluation card export
  • PlannedSecond consumer: SAT AI / claims verification