AI-model benchmark august 2026
AI model benchmark August 2026: three models score 25/25, local model Qwen 3.8 beats Fable for bug fixing
New models immediately join the leading group. At the same time, quality, cost and active execution time vary significantly.
We have expanded our benchmark for AI-assisted software development with new models and model versions. The results from 31 August 2026 make one thing particularly clear: a single overall model ranking is becoming less useful.
In full-stack software generation, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Terra all achieve the maximum score of 25/25. Nine models follow at 24/25.
The bug-fix test separates the models much more clearly. GPT-5.6 Sol leads with 68 out of 71.5 points. Claude Opus 5 follows at 60.5, while Grok 4.6, GPT-5.6 Terra and Qwen 3.8 27B FP8 finish close together at around 57 points.
These differences between tasks are particularly relevant for organizations that want to use AI systematically in software development. The question shifts from "which model is best?" to "which model performs well on this task, at what cost and within what timeframe?"
Two different software tests
The benchmark consists of two parts. The first test evaluates full-stack software generation in a Kanban use case. Each run consists of four tasks. The result can receive a maximum of 25 points across five evaluation dimensions:
scope fidelity
architecture
UX completeness
operational quality
reliability risk
Each dimension is worth a maximum of five points. The current chart contains 21 model entries.
The second test focuses on bug fixing in an existing codebase. It contains 29 seeded defects. The maximum weighted score is 71.5 points. Critical defects are worth 5 points, High 3, Medium 2, Low 1 and Trivial 0.5. A fix only counts when the corresponding verification step passes. One round was performed per model. Where available, we also record wall-clock time and run cost. Wall-clock time represents active building or fixing time, excluding periods when a run was idle.
The cost figures are not all directly comparable. Some subscription runs are represented using API-equivalent costs. GPT-5.6 costs in the bug-fix test, for example, are estimated based on token counts. Qwen 3.8 27B FP8 ran on self-hosted hardware and therefore has a $0 provider charge. Compute, hardware and management are of course not free. For several runs, no reproducible time or cost value is available.
Full-stack: the top of the ranking is approaching the ceiling

Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Terra each achieve the maximum score of 25/25.
Directly behind them is a notably broad group at 24/25: GLM 5.3, Grok 4.6, Kimi K3, Codex 5.5, GLM 5.2, Fable, GPT-5.6 Luna, Muse Spark 1.2 and Qwen 3.8 27B FP8. For model selection, the difference between 24 and 25 points is therefore less informative than the score alone might suggest. Cost and build time provide much more differentiation within this leading group. Claude Opus 5 achieved 25/25 in 3 hours 40 minutes for $35.87. Claude Opus 4.8 reached the same score in 2 hours 4 minutes for $24.74. GPT-5.6 Terra cost $17.58 and required 4 hours 2 minutes.
The quality score is identical, but the operational profiles of the runs differ considerably.
The results also vary within the 24/25 group. Grok 4.6 achieved 24 points in 1 hour 26 minutes for $19.52. Muse Spark 1.2 required 2 hours 1 minute for $18.10. Fable also scored 24/25, but its recorded run cost was $59.29. GPT-5.6 Luna reached the same score at approximately $6 in API-equivalent costs. Its active build time was 6 hours 44 minutes. Qwen 3.8 27B FP8 is particularly relevant from a self-hosted perspective. The model scored 24/25 with no provider charge, but the run took 10 hours 21 minutes. That is the longest recorded full-stack run in the chart.
Differences remain relevant below the leading group
Five models score 23/25: DeepSeek V4 Pro 0813, GPT-5.6 Sol, Composer 2.5, Grok 4.5 and Gemini 3.6 Flash.
The gap then grows. Gemini 3.7 Flash and MiniMax 2.7 score 20/25, DeepSeek V4 Flash 0731 reaches 16/25 and Qwen 3.6 Plus 12/25.
The cheapest run is therefore not automatically the most attractive. MiniMax 2.7 cost $1.11 in this test and Qwen 3.6 Plus $0.36, but their quality scores are clearly below those of the leading group.
The amount of additional human correction or review required by those lower scores was not measured. This benchmark therefore does not allow us to calculate total development cost per model.
Bug fixing creates much greater separation between models

While the full-stack results are tightly grouped at the top, bug-fix scores range from 27 to 68 points.
GPT-5.6 Sol leads with 68/71.5. Claude Opus 5 follows with 60.5/71.5. Grok 4.6 comes next with 57.5, GPT-5.6 Terra with 57 and Qwen 3.8 27B FP8 with 56.75. Kimi K3 scores 54.5 points. Fable and Grok 4.5 both reach 52, while Muse Spark 1.2 scores 50. Further down the ranking, DeepSeek V4 Pro 0813 scores 35.5, Qwen 3.6 Plus 31.5, DeepSeek V4 Flash 30.5, DeepSeek V4 Flash 0731 28.25 and MiniMax 2.7 27.
This test makes differences between models more visible than the current full-stack rubric. One reason is that every bug is predefined and a fix only counts when the verification step passes.
The same model can rank very differently depending on the task
GPT-5.6 Sol ranks first in bug fixing with 68/71.5, but scores 23/25 in the full-stack test. Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Terra all achieve 25/25 there.
Claude Opus 4.8 shows almost the opposite pattern. It reaches the maximum full-stack score, but scores 38/71.5 in bug fixing.
Claude Opus 5 performs strongly in both tests: 25/25 for full-stack and 60.5/71.5 for bug fixing.
GPT-5.6 Terra also combines high scores across both tasks, with 25/25 and 57/71.5.
These differences are a good reason to evaluate models per task. A model that performs particularly well at repairing existing code does not automatically rank highest when generating a complete application feature, and vice versa. Two tests are of course too limited to support general conclusions about software development as a whole.
Newcomers make the bug-fix leading group more interesting
Compared with our previous benchmark, several new model versions have been added.
In the previous full-stack benchmark, Claude Opus 4.8 was the only model at 25/25 and GPT-5.6 Terra scored 24/25. In the current benchmark, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Terra all achieve 25/25. In bug fixing, several previously tested scores remain visible, including GPT-5.6 Sol at 68/71.5, GPT-5.6 Terra at 57, Fable at 52 and Claude Opus 4.8 at 38. New entries in the current leading group include Claude Opus 5 at 60.5, Grok 4.6 at 57.5 and Qwen 3.8 27B FP8 at 56.75.
Grok 4.6 in particular combines three interesting values in this specific run: 57.5/71.5 points, 8 minutes of active fixing time and a $1.35 run cost. Claude Opus 5 scores three points higher at 60.5, with 35 minutes of active time and $14.54 in recorded costs. Qwen 3.8 27B FP8 reaches 56.75, almost the same scoring level as Grok 4.6 and Terra, but required 47 minutes and ran on self-hosted hardware.
These results make the models relevant candidates for further testing on an organization's own codebase. The data is too limited to create a general price-performance ranking. Comparisons between benchmark rounds also require caution. The charts do not contain all prompt, configuration and environment details required to treat every change as a fully controlled longitudinal comparison.
Quality, time and cost produce three different rankings
Sorting purely by quality produces a different result from sorting by cost or speed. This is immediately visible among the three full-stack models scoring 25/25. Their recorded costs range from $17.58 for GPT-5.6 Terra to $35.87 for Claude Opus 5. Claude Opus 4.8 is the fastest of the three at 2 hours 4 minutes, while Terra required 4 hours 2 minutes.
The relationship changes again in bug fixing. Grok 4.6 scores 57.5 points in 8 minutes for $1.35. GPT-5.6 Terra reaches 57 points at approximately $4.05. No reproducible wall-clock value is available for Terra in this run. GPT-5.6 Sol reaches 68 points at approximately $8.44. No reproducible timing value is available for that run either.
The most attractive combination depends on the workload. For a high-risk change, quality may matter more than a few dollars in inference cost. With thousands of small, easily testable changes, price may become more important. For interactive tasks, waiting time can become a practical bottleneck. Managing these three metrics separately is therefore more useful than combining them into one overall ranking.
Model routing becomes more relevant as task performance diverges
In a production environment, using one default model for every development problem becomes less attractive as the differences between tasks increase. A task can first be classified by type, risk and acceptance criteria. A suitable model can then be selected based on internal test results for quality, speed and cost.
This benchmark provides several concrete candidates.
For full-stack generation, Claude Opus 5, Claude Opus 4.8 and GPT-5.6 Terra lead with 25/25.
For the tested bug fixes, GPT-5.6 Sol achieves the highest score, followed by Claude Opus 5.
Grok 4.6 stands out in the bug-fix test for its combination of 57.5 points, 8 minutes of active time and a $1.35 run cost.
Qwen 3.8 27B FP8 is relevant for organizations evaluating self-hosted models. Its 24/25 full-stack score and 56.75/71.5 bug-fix score need to be weighed against its longer execution times and internal infrastructure costs.
Within the SiliconCode AI Software Factory, we use this type of measurement to evaluate model selection in combination with architecture, quality controls and the type of development task.
Limits of this benchmark
The bug-fix test consists of one round per model. The full-stack test uses four tasks per run and a qualitative composite score. These results allow us to identify relevant differences, but they do not provide statistical certainty for future runs. The model sets are also not identical across both tests. Codex 5.5, for example, appears in the full-stack chart but not in the current bug-fix chart.
Timing or cost data is missing for several runs. Provider charges, API-equivalent values and self-hosted compute also represent different types of costs. The benchmark does not measure how much human review was required after each run, how much output eventually reached production or the total infrastructure costs involved.
The results therefore answer a specific question: how did these particular models and versions perform on these two software tests?
Conclusion
The full-stack benchmark now has a broad leading group. Three models achieve 25/25 and another nine score 24/25. The current rubric therefore provides relatively little differentiation at the top.
The bug-fix test produces a clearer ranking. GPT-5.6 Sol leads with 68/71.5. Claude Opus 5 follows at 60.5, with Grok 4.6, GPT-5.6 Terra and Qwen 3.8 27B FP8 close behind.
Three points are particularly relevant for model selection:
Performance demonstrably differs by task.
Quality, cost and execution time produce different rankings.
New model versions can immediately become relevant shortlist candidates within a single benchmark round.
Running your own benchmark on representative software tasks therefore provides more useful evidence than relying on a single general public ranking.
Want to see how SiliconCode uses multiple AI models within a controlled software development process? Request a personal SiliconCode demo or compare these results with our previous AI model benchmark.