Matt Wolfe
July 31, 2026
TL;DR
Claude Opus 5 matches GPT-4.5 on benchmarks at lower cost but receives widespread criticism for poor reasoning quality; it excels uniquely at generating visually impressive game graphics with 3D.js.
“Opus 5 could be one of the first regressions in modern LLM history. It's rather verbose and scattered in thinking. Not pleasant at all for knowledge work.”
— User feedback
“It like never answers the actual question you ask it. Claude Opus 5 is clearly a downgrade from Opus 4.8. It feels nerfed, not upgraded.”
— User feedback
“The gameplay itself still really, really sucks. Like, I don't think my attacks are hitting. My character just kind of looks like a blob with a sword that's going the wrong direction.”
— Creator testing Opus 5 game generation
“I just have a hard time being a fan of it in general because it just means we're going to see so much more AI slop and like low quality slapped together content just thrown on the web.”
— Creator on AI avatar video tools
1. Claude Opus 5 Release and Benchmark Performance
Opus 5 launched Friday after the previous week's news video, matching GPT-4.5 on benchmarks (agentic coding, terminal coding, knowledge work) with a 5.6-point margin in some categories, but at significantly lower cost than GPT-4.5
2. Real-World Performance Issues and User Backlash
Despite benchmark success, Opus 5 receives extensive criticism from users for being verbose, scattered, and making regressive mistakes compared to Opus 4.8; described as having poor personality, treating every issue as P0 severity, and failing to answer actual questions asked
3. Opus 5's Strength: AAA-Quality Game Generation
Opus 5 uniquely excels at creating visually impressive 3D.js game graphics—notably better than GPT-4.5—with a specific prompt formula repeated on GitHub using the word 'utterly' multiple times; generates full games with sound in 10-12 hours from single prompts
4. Game Generation Testing and Comparison
Creator tested Opus 5 with Elden Ring-style fantasy game prompt, producing 11-12 hour generation with impressive mountain and tree graphics but broken gameplay (reversed controls, non-functioning attacks); GPT-4.5 created more playable but less visually impressive results
5. Google Earth AI Integration and Demonstrations
Google integrated Nano Banana directly into Google Earth, allowing free visualization of historical reconstructions (Pompeii 78 AD) and future scenarios; tested with Petco Park, predicting biodomme, plant-based buildings, and magic streaks in 100-year future
6. Meta AI Platform and Multi-App Integration
Meta AI can now manage calendar and Gmail, read inbox emails and draft replies, add calendar events—late to market compared to OpenAI and Google but expanding agentic capability to third-party apps
7. Buzz: Decentralized Slack Alternative with AI Agents
Jack Dorsey created Buzz, a free open-source Slack competitor where users can invite Claude, Codeex, and Grok models as team members; supports agent-to-agent feedback loops and collaboration—demonstrated with three agents competing to build HTML websites
8. Video and Avatar AI Tools Roundup
HeyGen created two-host video podcast generator from documents (2:40 video tested on Alpha Evolve paper); Mirage Avatar X costs $30/month for expressive AI avatars; Google released Omni video model free through August 4, 2026 (10 free videos in Gemini)
9. Physical Robotics and Google's Gemini Robotics ER2
Google released Gemini Robotics ER2 model powering robot brains for complex tasks—demonstrated unscrambling grapes, unscrewing lightbulbs, and tying trash bags; Trump administration banning new Chinese humanoid robots over espionage concerns