Skip to content

Problems.

Write the piece, drive the whole agent loop, or repair a repo by prompting a coding agent. Most are graded against a real model, because that is what you will ship against.

0 of 85 solved

Show me at , . 85 matches

  1. 16Score an Answer with an LLM Judgeimplement · liveagent · liveL07medium
  2. 17Ignore the Instructions Hidden in Your Documentsdebug · liveagent · liveL08hard
  3. 18Answer an MCP HandshakeimplementstdlibL06easy
  4. 19Serve a Tool Call over MCPimplementstdlibL06medium
  5. 20Never Reply to an MCP NotificationdebugstdlibL06medium
  6. 21Dispatch a Model's Tool CallimplementagentL04easy
  7. 22Refuse the Arguments a Model InventedimplementagentL04medium
  8. 23Stop Paying for the Same Answer Twiceimplement · liveagent · liveL08medium
  9. 24Measure Retrieval with Recall@k and MRRimplementnumpyL02medium
  10. 25Route a Question to the Right Toolimplement · liveagent · liveL04medium
  11. 26Make the Model Admit It Doesn't Knowimplement · liveagent · liveL03hard
  12. 27Score a Model Against a Golden Setimplement · liveagent · liveL07medium
  13. 28Answer Over a Context That Does Not Fitimplement · liveagent · liveL08hard
  14. 29Build a LangGraph With a Conditional Edgeimplement · livelanggraphL05medium
  15. 30Fix a LangGraph That Forgets Everythingdebug · livelanggraphL05hard
85 problems, showing 16 to 30