Failure modes, not vibes
Trajectories are clustered into named failure modes you can read, dispute and prioritise.
Agent post-training
We read your agent's trajectories, find where it fails, build the tasks that target those failures, and train against them. Held-out evals show whether the score really went up.
Approach
Average traffic is already handled. Gains come from the rare, hard cases, so we start from your real failures and make more of them.
Trajectories are clustered into named failure modes you can read, dispute and prioritise.
Each mode becomes new tasks with graded rubrics, reviewed by hand before anything trains on them.
Tasks are split into held-out evals and training tasks up front, so the reported gain is on cases the model never trained on.
Beyond weights, we add auditing skills to the agent harness, which raised performance further.
How it works
You share agent runs, with outcomes if you have them.
We identify and rank how the agent fails.
New tasks per failure mode, split into eval and train.
Our RL pipeline runs on our infrastructure. You do not manage GPUs.
You receive the eval report, the tasks and the improved agent.
Case study
We ran the full loop on a deployed agent that edits and audits financial workbooks: failure modes from real trajectories, new tasks, an eval/train split, RL, and a harness-side auditing skill.
Scores on our in-house auditing benchmark. The base model with the auditing skill alone scores 50.6%.
Fully managed
Data pipeline, task generation, evals, RL training and versioning are ours.
curl https://api.upshiftlabs.llc/v1/chat/completions \
-H "Authorization: Bearer $UPSHIFT_KEY" \
-d '{"model":"your-agent-v2","messages":[...]}'