To understand how a handpicked group of strong, cost-efficient models performs inside Foyflow, we ran six models through eight frozen legal assignments drawn from Harvey’s open-source Legal Agent Benchmark. The tasks covered drafting, review, analysis, and research across eight practice areas. Every model received the same matter documents, prompts, tools, 30-turn limit, and high-effort setting, and each result was graded against the same 377 criteria. This allowed us to test the complete legal-agent workflow: reasoning across documents, using tools, following instructions, and producing the requested Word or Excel deliverables.

Grok 4.5 achieved the highest overall score at 85.9% and was the only model to complete every requested deliverable. GLM 5.2 followed at 74.3%, with Kimi K3 at 72.1%, GPT-5.6 Luna at 66.8%, DeepSeek V4 Flash at 58.9%, and Sonnet 5 at 54.4%. One of the most useful findings was that lower scores did not always mean weaker legal reasoning. DeepSeek, for example, passed 82.2% of the criteria across the six tasks it completed, but two unfinished deliverables accounted for 69% of all its missed points. Our benchmark deliberately treats completion as part of quality: an agent that identifies the right issues but never produces the requested work product has not completed the assignment.

Looking at cost alongside quality changes the picture again. Grok delivered the strongest work, while GLM offered the best observed balance of quality and cost. Luna and DeepSeek were substantially cheaper, but their lower completion and criterion scores make the trade-off visible. This is why we view the benchmark as a tool for choosing models inside Foyflow, not as a universal model leaderboard. Our evaluations are intended not only to identify the highest-quality models, but also to track which models and solutions are suitable for fully local AI and cost-efficient use at scale. The study is still small and directional, but it highlights something important: for legal agents, the combination of model, harness, tools, and reliable delivery can matter more than raw model capability in isolation.






