Ravi Menon
Models & Benchmarks Editor
Ravi Menon tracks new AI model releases for AI just dropped. Every spec, context window, price and capability on a model page is sourced to the provider's own API, model card, or official docs and stamped with the date it was checked — vendor benchmark claims are labelled as vendor-reported, never restated as fact. When a model is preview-only or we haven't run it ourselves, the page says so instead of inventing a score.
Reviews by Ravi Menon

Nano Banana 2 Lite
Google's fastest and cheapest Gemini image model. I ran it: a text-heavy poster came back in 4.3 seconds with both lines of type rendered perfectly — the failure mode most image models still trip on.
Claude Sonnet 5
Anthropic's Sonnet-tier model — 'the best combination of speed and intelligence.' Near-frontier on coding and agentic work at roughly a third of Opus's price, with a 1M-token context and adaptive thinking on by default.
Claude Mythos
Claude Mythos is Anthropic's cybersecurity-focused, Mythos-class model — built to autonomously find and fix software vulnerabilities. Its rollout is restricted; a safeguarded public sibling, Fable 5, carries Mythos-class capability to the broader public.
GPT-5.6
OpenAI's next flagship after GPT-5 — announced June 26, 2026 as a preview and generally available since July 9, 2026 in three tiers (Sol, Terra, Luna) with published per-token pricing.
Qwen-AgentWorld
Alibaba's Qwen team open-sources AgentWorld — a language world model that simulates seven agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) and is trained to predict how each environment responds to an action. Released June 24, 2026 under Apache 2.0.
GLM-5.2
Zhipu's open-weight coding flagship: a 753B mixture-of-experts model with a 1M-token context and MIT weights, claiming to edge past GPT-5.5 on coding benchmarks.
VibeThinker-3B
A 3-billion-parameter open reasoning model that claims to match systems hundreds of times its size on math and code — and has the AI world arguing about whether the benchmarks are real.
Gemma 4
Google DeepMind's fourth-generation open-weight model family — five sizes from 2B to 31B, Apache 2.0 licensed, with the 12B Unified variant accepting text, image, audio, and video in a single encoder-free architecture.
MiniMax M3
MiniMax's third-generation flagship — M3 is an open-weight 428B-parameter MoE (~23B active per token) using MiniMax Sparse Attention (MSA), with a 1M-token context window and native multimodality (text, image, and video input). It's positioned as the first open-weight model to combine frontier coding, a 1M context, and native multimodality — and can even operate a desktop computer.
Claude Opus 4.8
Anthropic's most capable Opus-tier model, built for complex agentic coding and enterprise work — a 1M-token context, adaptive thinking, and 128K max output, hosted only via the Claude API.

Kling 3.0
Kuaishou's flagship hosted video model — cinematic text- and image-to-video with native audio. I ran it: clearly a tier above the open models, with the usual fast-motion artifacts.

Pyramid Flow
An open-source, MIT-licensed text- and image-to-video model that makes 768p, 24fps, 10-second clips — and runs on a single consumer GPU with offloading.
See how we review for our sourcing and rating standards.