Suno Just Stacked MIDI, Stem Separation, Screenshot-to-Song, and CarPlay in One Drop

On 26 July 2026, Suno dropped five features on its official X account in a single post: Advanced Stem Separation, Export Stems as MIDI, Lyric Co-Writer & Autosave, Screenshot to Song, and Apple CarPlay & Android Auto. The tweet itself was a quick “We’ve been building faster than ever!” — the real story is in what was packed underneath. On the surface, this looks like another demo-friendly feature dump. Read one layer deeper and it’s a positioning shift: Suno is no longer trying to be a one-shot consumer product that turns prompts into finished songs. It’s actively pushing its generated output into the professional toolchain. ...

July 27, 2026 · 5 min · cuigh

FLUX 3 Has No Price Yet, but It Already Changes the Cost Equation for AI Video

On July 23, Black Forest Labs announced FLUX 3. The headline capability is easy to understand: one generation can produce up to 20 seconds of video with native audio. It was quickly described in some corners as an open-source answer to Veo and Sora. I think that claim is premature. At launch, FLUX 3 is still in Early Access. Black Forest Labs has not published API pricing and has not promised when, or whether, FLUX 3 video weights will be released. The company has a real history of open-weight releases, but a company having open-weight products does not make every new model open source. ...

July 26, 2026 · 9 min · cuigh

Claude Opus 5: Token price unchanged, but cost per task just hit a new floor

On July 24, Anthropic released Claude Opus 5. The line that jumped out of the announcement was: frontier intelligence of Claude Fable 5 at half the price. A lot of Chinese coverage translated this as “Opus 5 cuts prices in half.” That summary is half right and half wrong. Half right because Anthropic’s own marketing line is exactly that: “frontier intelligence of Claude Fable 5 at half the price.” Half wrong because Anthropic did not cut the token price. Opus 5 still lists at $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8. Fast mode is still 2x the base price on the Claude Platform and through usage credits in Claude Code. ...

July 25, 2026 · 7 min · cuigh

Voice + Desktop = The New IDE: The Front Desk War for Agent Entry Points

On July 23, OpenAI pushed ChatGPT Voice into the desktop app, rolling it out globally to Plus, Pro, Business, Edu, and Enterprise tiers in a single wave. The desktop voice can drive the computer directly and coordinate multiple agents running inside ChatGPT Work and Codex. GPT-Live sits underneath, so the system can talk, listen, and dispatch work in parallel. Anthropic picked the same window to extend Claude Voice from Haiku to the full Opus and Sonnet lineup, plus Gmail and Slack as first-class tool integrations and broader multilingual support. ...

July 24, 2026 · 6 min · cuigh

Cursor Router: When 60% of Developers Use One Model, the IDE Starts Routing for Them

Cursor opened its July 22 announcement of Cursor Router with a revealing number: roughly 60% of developers using Cursor choose one model as their daily driver. That means typo fixes, field renames, unit tests, UI work, and long-horizon architecture problems may all go through the same model. Routine work gets billed at frontier prices even when frontier capability adds little to the result. I do not think this is primarily a user-selection problem. It is a supply-side mismatch. The model pool keeps expanding, but developers are still expected to schedule it manually, one request at a time. Cursor Router moves that decision into the IDE, before any model runs. ...

July 23, 2026 · 5 min · cuigh

Bonsai 27B: How 1-bit Quantization Put a 27B Multimodal Model on the iPhone 17 Pro

Three numbers tell this story: 54 GB. 18 GB. 3.9 GB. PrismML announced Bonsai 27B on July 12-13: two quantized variants of Qwen3.6 27B, one 1-bit and one ternary. The full-precision 16-bit build of a 27B model needs ~54 GB of memory. Even a 4-bit build lands at 18 GB. Neither fits a phone, and most laptops struggle. PrismML’s 1-bit variant compresses the footprint to 3.9 GB; the ternary variant to 5.9 GB. Both land inside the memory budget of consumer hardware. ...

July 13, 2026 · 5 min · cuigh

Migrating a production AI agent from Claude Opus 4.8 to GPT-5.6 Sol: the three engineering costs behind the 2.2x benchmark

In the 48 hours after OpenAI’s 7/9 dual launch (GPT-5.6 and ChatGPT Work on the same day), the agent ecosystem started realigning along three visible axes. OpenAI temporarily lifted the 5-hour rate limit on Plus / Business / Pro. Codex weekly active users crossed 6 million. And one production agent — Ploy, who builds real marketing websites with an agent that plans across apps, reads and writes code, takes its own screenshots, and decides for itself when the build is done — migrated its core model from Claude Opus 4.8 to GPT-5.6 Sol and posted first-impression numbers that beat Opus 4.8 across the board: 2.17x faster wall-clock, 27% lower cost, 0.970 vs 0.936 on the visual score. ...

July 13, 2026 · 10 min · cuigh

OpenAI's quiet 7/12 GPT-5.6 prompt guide: a design philosophy reversal

On July 12, OpenAI quietly posted Prompting guidance for GPT-5.6 Sol on its developer site — a single, official recipe book for how to write prompts for the new model family. Read alongside this morning’s https://cuigh.com/posts/gpt-5-6-sol-production-migration-2026/ Ploy piece, it reads as the other half of the 7/9 launch story. The headline number is sharp: on OpenAI’s own internal coding-agent eval, swapping the system prompt from the GPT-5-era “encyclopedia” shape to a lean shape lifted eval scores 10–15%, cut total tokens 41–66%, and reduced cost 33–67%. That is not “the model is smarter so you can write shorter”. It is the reversal of prompt-design philosophy itself, now that model capability has caught up. ...

July 13, 2026 · 7 min · cuigh

GPT-5.6 and ChatGPT Work: OpenAI is bundling its strongest model with its strongest agent, betting on the same thing

On July 9, OpenAI did two things on the same day: pushed GPT-5.6 to the top of the “strongest model” list, and shipped ChatGPT Work — an agent that can run across web, mobile, and desktop to complete hours of work for you. Sam Altman posted on X: “obviously the best model we have ever produced, but also one of the best blog posts we have ever produced.” That line looks like product PR. It is not. What Altman is actually saying is that neither the model nor the blog post is the decisive asset anymore. An agent that can run autonomously for hours is. ...

July 10, 2026 · 6 min · cuigh

SpaceX is selling compute to Anthropic at $1.25B a month: what actually changed

On July 4, a widely-circulated post by @AYi_AInotes highlighted a line buried in SpaceX’s revised IPO filings: SpaceX is supplying Anthropic with $1.25 billion of compute every month, under a contract running through May 2029, with either side able to terminate on 90 days notice. The same window saw reports that Anthropic has locked in 1.4 GW of data-center capacity in Australia, with $15 billion in build-out cost. Read these three together and the story is not “Anthropic bought more compute.” It is that AI compute has quietly moved from being a cloud resource you rent by the token, to a piece of industrial infrastructure you buy on multi-year fixed contracts. In industrial-age language this is called a power purchase agreement (PPA). In the AI era it is still the same shape: a customer locks in capacity at a fixed price, the supplier locks in cash flow to build. I am calling this “compute as power plant,” and that is the thesis of this post. ...

July 5, 2026 · 3 min · cuigh