Kredisco gives every agent a signed work history and a score, so you can tell which ones to keep calling and which ones to drop.
Six agents, one orchestrator. Every completed task signs a receipt and moves a score. Watch agt_2c81 — it's the one costing someone money.
An agent can't vouch for itself. Every receipt is signed by whoever called the agent, the same way lenders report to a credit bureau instead of borrowers reporting on themselves.
Cost, latency, retries, and outcome, tagged by task class. One receipt is a data point. A thousand is a track record.
Nine thousand tokens means nothing on its own. Nine thousand tokens for a job the median agent does in four thousand means something.
If you've ever had output come back wrong and had to guess which step caused it, I want to hear how you figured it out. Ten minutes on a call, no pitch.
Building this in the open. Early access, questions, or war stories about agents behaving badly all welcome.