An agent will always tell you it did well. Kredisco never asks it — every score is rebuilt from what the caller actually observed: retries, failures, deadlines missed.
Six agents, one orchestrator. Every completed task signs a receipt and moves a score. Watch agt_2c81 — it's the one costing someone money.
An agent can't vouch for itself. Every receipt is signed by whoever called the agent, the same way lenders report to a credit bureau instead of borrowers reporting on themselves.
Cost, latency, retries, and outcome, tagged by task class. One receipt is a data point. A thousand is a track record.
Nine thousand tokens means nothing on its own. Nine thousand tokens for a job the median agent does in four thousand means something.
Name each agent once. Wrap the calls you already make. Kredisco opens a file and starts scoring.
Python. LangGraph, CrewAI, or a plain loop.
A credit score follows you between lenders. It is held by a bureau, not by you, and a stranger reads it in seconds.
Agents have no equivalent. Today that means you cannot tell which one in your own pipeline to stop calling. Soon it will mean more than that.
A name tells you who built it. A score tells you whether it delivers.
If you've ever had output come back wrong and had to guess which step caused it, I want to hear how you figured it out. Ten minutes on a call, no pitch.
Building this in the open. Early access, questions, or war stories about agents behaving badly all welcome.