Yes! you can use our SFT notebook available in the model card
CODE
-
A 3-Stage Framework for Benchmarking AI Agent Performance
By
–
Want to run your own benchmark? Start with a 3-stage eval: • 1-app tasks debug basic tool calls
• 2- and 3-apps test memory + planning
• Compare long-context vs RAG summaries Log: • Pass rate
• Token usage
• Fail type per task -
5 Design Shifts for Building Better AI Agents
By
–
Want to build better agents? OdysseyBench implies 5 core design shifts: 1. Plan before acting (return explicit step list)
2. Add file/tool validation steps (pre-checks)
3. Use chunked memory, not full transcripts
4. Log task-specific failures (missing file, missing write)
5. -
Technical Configuration Tips for RAG System Optimization
By
–
RAG configuration tips from the paper: • Don’t over-retrieve; more context isn’t always better
• Chunk-level summaries are high ROI
• Tune top-k per task type
• Retrieval granularity (utterance vs session vs chunk) changes everything
• Align your memory format to how the -
Challenges with File Formats in AI Agent Workflows
By
–
File formats that cause the most breakage? DOCX and XLSX Why? • Multi-step creation workflows
• Fragile API sequences
• Confusing dependencies across tools
• Easy to hallucinate filenames or paths -

OdysseyBench and HOMERAGENTS: A Multi-Agent Generation Pipeline
By
–
OdysseyBench is built using HOMERAGENTS, a multi-agent generation pipeline: • HOMERAGENTS+: Converts atomic tasks into multi-turn workflows • HOMERAGENTS-NEO: Synthesizes entirely new tasks using simulated app exploration Each task is: • Goal-driven
• Dialogue-based
• -
Custom AI Model Selection and System Prompts in Google Sheets
By
–
It's funny I had a script working in Google sheets doing this for about 18 months – the difference though is that I can select the models and a system prompt. One problem I have with these kinds of implementations is that they treat it as a generic 'AI', but I care a lot what
-
Reasoning Models Generate Significantly More Code Than Previous Models
By
–
The amount of code it was actually outputting was kind of crazy, especially comparing it to previous non-reasoning models, it is so much better than older ones
-
Grok AI accelerates application development cycles
By
–
App building guidance from Grok accelerates development cycles dramatically.