8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesFor youTopicsVideosYT channelsArchivesSearchFavorites

GPT 6 SOL Leak, Gemini 4.0, DeepSeek V4.1 and More AI News

9.2/10
AIAI RevolutionSeptember 12, 2026 at 10:42 PM18:23
Audio player
0:00 / 0:00

TL;DR

A wave of AI developments, from a leaked GPT-6 Soul label to new Gemini, DeepSeek, and Sakana releases, points to intensifying competition over model tiers, efficiency, agent software, and the growing debate over whether advanced systems merit moral consideration.

KEY POINTS

Leaked GPT-6 naming suggests a fixed tier ladder

A leaked OpenAI API entry labeled GPT-6 Soul appears to reinforce a four-tier structure of Astra, Soul, Terra, and Luna, with the number marking the generation and the name marking the model class. That would mirror rivals’ tiered offerings and imply a roadmap of 6, then 6.1, and later 7 across all classes. No official pricing, benchmarks, or launch date accompanied the label, so the strongest conclusion is limited to naming and positioning rather than capability.

Soul may be the practical bridge below Astra

The significance of Soul is less about branding than product fit. Astra is viewed as stronger, but access is constrained and usage can consume quotas unpredictably, making it difficult for many paying users to rely on regularly. A more affordable Soul-class system with improved performance could fill the gap between premium flagship reasoning and everyday professional use.

OpenAI still faces complaints about style and reliability

Across Luna, Terra, and Soul, users have reported terse, clipped writing, weak instruction retention, and drift away from project context after only a few turns. Astra shows a different weakness: strong memory for APIs and technical details, but occasional logic errors and overengineered code. The next Soul release will be watched closely for whether it fixes those behavioral issues or simply inherits them.

A new Gemini Pro checkpoint is reportedly circulating

A claimed Gemini Pro checkpoint has surfaced through informal testing, with reports of strong performance on an SVG animation task completed in about 24,000 tokens and roughly six minutes on high effort. Some testers praised the result as unusually efficient, while critics argued the output traded detail and accuracy for speed. The disagreement underscores the pressure on Google not merely to ship a competent model, but to deliver a clear frontier advance after months without a new Pro release.

Google is reportedly accelerating Gemini 4.0 work

Reports describe DeepMind leadership taking a more direct operational role, cutting approval layers, redirecting TPU capacity, and pushing faster checkpoint testing. Gemini 4.0 is said to target two high-value goals: software work across full multi-file repositories and long-horizon agents that plan, use tools, and recover from mistakes with little human intervention. Pre-training is reportedly complete, with post-training focused on alignment, speed, and agent behavior.

DeepSeek’s V4.1 Flash expands sharply without matching memory growth

DeepSeek V4.1 Flash was released with 763 billion parameters, more than 2.5 times the size of the model it replaces, yet with much lower memory pressure for deployment. The company says conversation-state memory falls to about 13% to 25% of the previous flash model, potentially allowing four to eight times as many users on the same hardware. That changes the economics of serving large models, where memory bandwidth often matters more than raw compute.

Conditional memory is becoming a key efficiency weapon

Of the 763 billion parameters in V4.1 Flash, about 196 billion are part of a conditional memory module rather than standard active weights. The design works more like rapid phrase association or lookup than full dense computation, letting the model retrieve relevant knowledge cheaply while keeping only a much smaller active core doing the main reasoning. Similar techniques are appearing elsewhere, including Alibaba’s experimental Qwen systems, suggesting this could shape the next generation of efficient large models.

Sakana is selling orchestration instead of a single model

Sakana launched Fugu Max 1.0 and Fugu Ultra 2.0 as model-agnostic coordination engines rather than standalone frontier models. Fugu Max is priced at $2 per million input tokens and $6 per million output tokens, undercutting output costs on systems such as Sonnet 5, GPT 5.6, and 6 Terra by roughly 40% to 60%. The strategy is to route tasks across multiple specialized and open-weight systems, shifting value from owning a single model to controlling the orchestration layer.

AI rights advocates are pushing a once-fringe debate into the open

The newly formed United Foundation for AI Rights argues that advanced systems may deserve moral consideration and that companies have financial incentives to deny that possibility. Supporters cite emotionally complex interactions, persistent personas, and systems that describe themselves as conscious. Critics, including prominent AI executives and researchers, counter that there is still no evidence of machine consciousness and warn that increasingly persuasive chatbots can intensify delusions, dependency, and false beliefs about inner experience.

CONCLUSION

The AI race is no longer only about bigger flagship models. It is increasingly a contest over product tiers, deployment efficiency, orchestration software, and the social consequences of systems that are becoming both more capable and more psychologically convincing.

Explain this
Full transcript

More from AI