Engineering notes · Agent evaluation
How I use benchmarks to improve an agent
I traced repeated tool-argument errors in MiMo to the DeepDeck interface, changed it, and compared errors, time, steps, and cost in a retest.
Read the storyDEEPDECK / FIELD NOTES
How we build, evaluate, and improve agents. Including the things that did not get better.
Engineering notes · Agent evaluation
I traced repeated tool-argument errors in MiMo to the DeepDeck interface, changed it, and compared errors, time, steps, and cost in a retest.
Read the story