Models No, Sakana Fugu Is Not a Model That Beats GPT and Claude Fugu is a router that calls GPT, Claude, and Gemini, not a rival to them. Its own report shows strong scores. It also shows the one comparison driving the hype cannot be checked by anyone outside Sakana. Last updated: 25 July 2026 The short version Fugu is an orchestrator: a small model that reads your request and routes it to a pool of frontier models (Claude Opus 4.8, Gemini 3.1 Pro, GPT-5.5), then returns one answer. It cannot be "better than" GPT or Claude in the way the hype says, because it works by hiring them. Every Fugu score in the report is reported by Sakana itself. No independent lab has verified them. claim The headline claim, that Fugu matches Anthropic's Fable 5 and Mythos, compares against two models that are not in Fugu's pool and are not publicly accessible, so no outsider can reproduce it. The real, narrower result: on Sakana's own numbers, Fugu beats the best single model in its pool on most benchmarks. Disclosure: Fugu orchestrates and is benchmarked against Anthropic's Claude Opus 4.8, and its headline claim compares against Anthropic's Fable 5 and Mythos. This newsroom uses Anthropic's Claude in production. What Fugu actually is Sakana Fugu is not a new frontier model. By Sakana's own description it is a family of orchestrator models: small language models trained to read a query, decide which frontier model in a pool is best suited to it, delegate the work, and synthesize a single answer. You call one endpoint. Behind it, Fugu is routing to Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5. There are two public variants, a fast one that picks a single model per task and a Ultra version that coordinates up to three, plus a cybersecurity-specific one. That design is the first thing the hype gets wrong. A claim that "Fugu beats GPT-5.5 and Claude" treats Fugu as a competitor to those models. It is not one. It is a layer on top of them. A dispatcher that assigns work to the best specialist on its team can outperform any one specialist working alone, but it is not a better engineer than the engineers it assigns. The correct claim is narrower and Sakana states it plainly: Fugu aims to beat the best single model in its pool, on a given task, by always routing to the right one. What the numbers say, and who reported them On Sakana's benchmark table, Fugu-Ultra posts the top score on most of the eleven benchmarks tested, including agentic coding, scientific reasoning, and multidisciplinary reasoning. On the two agentic-coding tests the margin over the next model runs about 5 to 6 percent, which Sakana likens to a full generational upgrade. claim Two things temper that before anyone repeats it as settled. First, the wins are not a clean sweep: GPT-5.5 still leads on long-context retrieval and on the MRCRv2 needle-in-a-haystack test, and several of Fugu's "wins" (on SciCode, on a banking dialog test, on GPQA where it ties) sit within the margin you would expect from noise. Second, and more important, all of Fugu's scores are produced by Sakana. The baseline scores are the model makers' own reported figures. No independent third party has benchmarked Fugu. That does not make the numbers wrong. It makes them a vendor's self-report, which is where every AI benchmark claim starts and not where it should end. Benchmark (as of Jun 2026) Fugu-Ultra Opus 4.8 Gemini 3.1 GPT-5.5 SWE Bench Pro 73.7 69.2 54.2 58.6 Terminal Bench 2.1 82.1 74.6 70.3 78.2 Humanity's Last Exam 50.0 49.8 44.4 41.4 GPQA Diamond 95.5 92.0 94.3 93.6 MRCRv2 (long context) 93.6 87.9 84.9 94.8 Selected rows from Sakana's Table 1. Fugu scores are Sakana-reported; baselines are provider-reported. GPT-5.5 still leads MRCRv2. The claim that does not hold up as science The line doing the most work in the hype is that Fugu stands "shoulder-to-shoulder" with Anthropic's Fable 5 and Mythos Preview, delivering frontier capability without the export controls those models carry. This is the claim to be most careful with, and the report itself explains why. In the caption of its own headline figure, Sakana notes that neither Fable 5 nor Mythos is in Fugu's agent pool, because they are not publicly accessible. Sit with what that means. A benchmark comparison is only meaningful if someone other than the party making the claim can run it and get the same result. Here, the two models Fugu is measured against cannot be called by Fugu, cannot be accessed by the public, and were scored using whatever figures were available to report. There is no way for an independent researcher to reproduce the comparison, because the comparison requires access to models that are locked. A claim you cannot test is not a benchmark result. It is an assertion wearing the costume of one. The "without export controls" framing is real and interesting as a business pitch, but it does not convert an unverifiable comparison into a verified one. To be fair to Sakana, it is not hiding any of this. It states that the scores other than Fugu's are provider-reported, and it states outright that the Anthropic models are not public. The problem is not concealment. The problem is that the most repeated claim rests on the least checkable evidence, and repetition has quietly promoted "Sakana says it matches two locked models, by its own account" into "Fugu matches the frontier." Those are not the same sentence. So where is the independent test? As of 25 July 2026, there is not one. Every serious write-up of Fugu says the same thing: the numbers are Sakana's own, and no third-party lab has reproduced them. They have not appeared on independent leaderboards. That is the crux. Saying you are the best carries weight only after someone who is not you runs the test. A score you generated, in a harness you configured, against numbers you selected, is the assertion, not the verification. The company's cybersecurity variant, Fugu-Cyber, shows how far self-reported scores can drift from an outside measurement. Sakana reports 86.9 percent on CyberGym, a benchmark of real-world vulnerability exploitation. claim But CyberGym's own creators, presenting the benchmark earlier in 2026, found that the best model combinations they tested cleared only about 20 percent. Sakana's self-reported figure is more than four times what the benchmark's authors measured for the strongest setups. That gap might reflect a genuine leap, or a different and undisclosed testing setup. There is no way to tell without an independent run, and there has not been one. The score is the entire quantitative case for the release, and it rests on Sakana's word. One clarification in fairness. For the locked Anthropic models, Sakana did not fabricate numbers; it pulled published scores from public leaderboards and services where Anthropic's own figures were unavailable. But lifting a competitor's leaderboard number and placing it beside your own self-run score is still not a head-to-head test. The two scores come from different setups, different harnesses, and different hands. Putting them in one table does not make them comparable. It makes them adjacent. What is actually worth watching Strip out the unverifiable comparison and a real result remains. If a trained router can take three frontier models anyone can already call and, by choosing the right one for each step, beat the best of them on most tasks, that is a genuine finding about orchestration as a way to get more out of existing models without training a bigger one. The in-pool comparisons are the defensible part: they use each model's own reported scores at matched reasoning effort. The idea deserves independent testing, on public models, by someone other than Sakana. Until that happens, the honest summary is this. Fugu is a clever router with strong self-reported numbers and one headline claim that cannot yet be checked. That is worth attention. It is not yet worth the word "beats." Sources Sakana AI, "Sakana Fugu Technical Report," arXiv:2606.21228, 19 June 2026 (orchestrator design, Table 1 benchmark scores, self-report and pool-access disclosures in the Figure 1 caption). Sakana AI, "Sakana Fugu" product page (variants, orchestration description, "without the risk of export controls" framing). TechTimes, "Sakana AI Fugu-Cyber Claims 86.9% Vulnerability Score; Benchmark Methodology Not Disclosed," 22 July 2026 (CyberGym self-report versus benchmark authors' ~20 percent finding; no independent reproduction at launch). DataCamp, "Sakana Fugu: Features, Benchmarks, and How It Works," 24 June 2026 (numbers Sakana-reported, not independently reproduced by third-party labs). Last updated: 25 July 2026. Figures labeled claim are operator-stated and not independently audited. All Fugu benchmark scores in this piece are reported by Sakana AI; no independent benchmark of Fugu exists as of this date.