- Planned Integration: An upcoming version of Sakana AI’s Fugu agent orchestrator is planned to include Nvidia Nemotron specialists, although Sakana AI has not disclosed an availability date.
- Model Routing: Fugu selects models for subtasks and combines their outputs, with Nemotron intended for coding, tool calling, and instruction following.
- Missing Evidence: Sakana AI has not named the Nemotron variant or published benchmark results for the planned pairing.
- Evaluation Plan: Nvidia will provide technical guidance, while both companies plan to monitor routing, synthesis, delay, cost, and reliability after deployment.
Sakana AI plans to add Nvidia’s open Nemotron models as specialists in a future release of its Fugu AI multi-agent orchestration platform. Fugu routes work among multiple artificial intelligence models, and Nemotron will handle coding, tool calling, and instruction following alongside its frontier models rather than replace them.
Fugu gives developers one API, selects models for individual subtasks, and combines their work into one response. Task-specific routing matches each job with a suitable model instead of sending every request to the same system. Sakana AI has not identified the Nemotron variant, set an availability date, or supplied benchmark results for the pairing, so its practical advantage remains untested.
How Fugu Will Divide the Work
Fugu’s routing layer decides which underlying model or agent should receive a request. A coding problem could go to a software specialist, while tool calling lets another model invoke external software or services. The planned expansion will test collective-intelligence architectures designed to balance accuracy, performance, and cost across open and proprietary systems.
Developers will continue sending requests through one interface while Fugu changes the underlying mix of specialists, avoiding an application rebuild around each new provider. Routing mistakes can erase that benefit because a misplaced task adds a call without improving the answer, while invoking several agents can increase latency and expense. Fugu must also resolve conflicts without discarding the detail that made delegation worthwhile.
Nemotron’s open format gives Sakana AI more control through open weights, datasets, and recipes, including downloadable parameters that organizations can customize and run on controlled infrastructure. Open weights still require hosting, representative evaluation sets, security controls, and tested updates, while a model change could alter response quality even when the application’s API remains the same. Nvidia plans to provide technical guidance, and both companies intend to monitor Nemotron after deployment by separating component quality from routing and synthesis performance.
The Benchmark Is Still Ahead
Nvidia includes Fugu just as one among Japanese AI projects using Nemotron. Other projects focus on industry-specific or locally adapted AI, while Sakana AI is applying the models to coordination across a pool. Adding a capable component and selecting it correctly remain separate engineering problems.
Nvidia previously introduced the Nemotron 3 family for agentic AI workloads and published Nemotron 3 Nano throughput data. Sakana AI has not identified the future Fugu specialist as that configuration.
Nemotron 3 Ultra has roughly 550 billion total parameters and 55 billion active parameters. Earlier Fugu benchmarking matched or outperformed older AI models like GPT-4 and Claude 3.5 on several reasoning and coding tests.
Nvidia says that Nemotron 3 Ultra delivered 5.9x, 4.8x, and 1.6x higher inference throughput than GLM-5.1-754B-A40B, Kimi-K2.6-1T-A32B, and Qwen-3.5-397B-17B, respectively. Nvidia used 8K input and 64K output for that test, which did not measure Fugu’s routing layer.
Nvidia’s throughput test measured one model under fixed input and output conditions, while the earlier Fugu result covered a different model pool. A deployed Fugu release must instead show whether its router chooses Nemotron for suitable tasks and combines the specialist’s output accurately.
Open models like Nemotron have already handled nearly a third of AI requests measured on Vercel’s AI Gateway in June, while closed models occupied a higher-cost premium layer. By July, model repository Hugging Face hosted almost three million public models and one million public datasets. That breadth gives orchestration systems more specialist candidates, but it also expands the selection and evaluation burden without establishing how this pairing will perform.


