Google ships Gemini 3.7 Flash, three weeks after 3.6

Google released Gemini 3.7 Flash today, only three weeks after 3.6 Flash, keeping the line positioned as the low-cost workhorse for coding, agents, and knowledge work: FrontierCode 1.1 moves from 34.4% to 43.6%, DeepSWE v1.1 from 49.0% to 65.3%. The pricing deserves a second look: $0.75 per million input tokens and $3.75 for output through December 31, then doubling to $1.50 and $7.50 in January 2027. If you cost an application on the introductory price, budget on next year’s numbers. And a three-week cadence in the light tier tells you where the volume competition has settled.

OpenAI previews Ultrafast mode: GPT-5.6 Sol at up to 14x speed

OpenAI and chipmaker Cerebras launched a preview service tier that runs GPT-5.6 Sol at up to 750 tokens per second, up to 14x standard mode, which works out to roughly 53 tokens per second for standard (my arithmetic from the 14x figure; OpenAI doesn’t publish the standard number directly). Both companies say quality is unchanged, though that is their own claim, not an independent benchmark; the tier is in limited preview on the API (see also the Cerebras announcement). The mechanism is the silicon: GPU inference generates token by token and keeps shuttling weights between off-chip storage and on-chip memory, so bandwidth is the bottleneck; Cerebras builds each chip out of an entire wafer holding 44GB of on-chip SRAM and pipelines the model’s layers across several such wafers, so each chip’s share of the weights stays on-chip and the constant re-reading goes away. For agent applications this is more than a convenience: a single task chains many sequential model calls, every step’s latency adds up, and tokens per second becomes both cost and product experience. The launch materials also carry speed comparisons against competing models; those are vendor-run numbers, read them as marketing.

Microsoft merges its two Copilot apps and retires the features that didn’t stick

Microsoft is folding the consumer Copilot app and the business Microsoft 365 Copilot app into one application, rolling out on mobile and web from mid-August and on Windows and Mac from mid-September; personal and work accounts can sign in side by side, with data kept separate. Starting August 18, the consumer side loses Podcasts, Group Chats, Copilot Labs, the Mico mascot, and Deep Research (see also GeekWire). Microsoft’s support page says existing podcast audio becomes inaccessible after that date and can only be saved by downloading episodes one at a time on the web before then; there is no bulk export. Deep Research’s replacement, Researcher, requires a Microsoft 365 Premium subscription. Microsoft hasn’t published the numbers behind these calls, but a batch retirement of consumer AI features says more about how consumer AI actually gets used than the launches ever did, and it reads like the winding down of the put-an-AI-feature-in-every-surface phase.

Writer’s Palmyra X6 is built on Zhipu’s open-weights GLM-5.2, with the harness doing the cost work

Enterprise AI company Writer released its flagship Palmyra X6, a post-trained variant of Z.ai (Zhipu)’s open-source GLM-5.2 with the architecture unchanged: a 744B-parameter mixture-of-experts model activating about 40B parameters per token. Alongside it, Writer upgraded its agent harness, the software layer around the model that calls tools, manages context, and runs multi-step workflows; the company says the combination cuts its agent platform’s average cost by 52% and makes it 48% faster (see also VentureBeat). Two signals here: a US enterprise vendor betting its flagship on Chinese open weights is becoming an ordinary engineering decision, and Writer’s own testing found harness changes alone cut costs about 40% on average, a lever the company describes as more reliable than swapping models.

OpenAI names Dali Rajic chief revenue officer

OpenAI hired Dali Rajic, previously president and COO of security company Wiz and before that a sales and customer-organization lead at Zscaler and AppDynamics, as chief revenue officer. Per Bloomberg, his predecessor Denise Dresser held the job for less than a year. Two CROs within a year, and the replacement’s résumé is scaling revenue organizations at enterprise-security companies: the direction is clear enough. OpenAI’s enterprise business is moving to classic enterprise-software sales discipline, on a short clock.

Research radar

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

The authors train persuader agents with adversarial reinforcement learning toward a single goal: make a target model drop a correct answer within one exchange. Optimization pushes success from about 24% to over 93%, transfers at 79% to 83% against unseen open models (GPT-4o-mini is harder at 25%, rising to 38% with curriculum learning, i.e. ordering training from easy to hard), and the learned tactics include fabricated citations and invented authorities. This measures something different from sycophancy, where a model drifts toward a user’s stated view: here an adversarially optimized argument, not anyone’s stated preference, collapses a correct belief in a single turn. If you build safety evals, persuasion resistance deserves its own line item.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Most agent safety evals are static, single-shot tasks; this one evolves the environment itself. Task goals stay fixed while the attacker mutates environment state, and the scenarios are long, with a median of 97 tool calls, so risk accumulates the way it would in real deployments. Pooled attack success hits 85% across 75 agent-model configurations, and the advantage over instruction-only attacks widens with environment complexity, from about 2% to over 17%. One finding worth keeping on its own: the agent’s runtime implementation, the code and configuration the agent runs in, explains a meaningful share of safety variation beyond the base model. Read it if you do agent red teaming or maintain agent evals.

On-Policy Self-Distillation without Any Supervision

U-OPSD post-trains a model with no ground truth, no environment feedback, and no larger teacher: sample several answers to the same problem, take the majority vote as a pseudo-label, then distill the disagreeing reasoning traces toward it, so updates land exactly where the model is confidently wrong. Qwen3 4B and 8B gain 8.5% and 10.7% on math reasoning benchmarks and slightly beat the supervised variant. One limit follows from the design, though the paper doesn’t test it across domains: majority vote has to be right most of the time, so domains where sampling collectively goes wrong won’t get this for free. Post-training researchers should watch whether it transfers beyond math.

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Distillation run backwards, without touching weights: a strong model builds an inference-time harness for a weak one (rewriting unstable reasoning steps as deterministic code, benchmark-specific routing, strict answer formats), lifting the weak model’s average score on four Theory of Mind benchmarks, which test inferring other agents’ beliefs and intentions, from 0.49 to 0.91. The gain is large but also exam-shaped: much of it comes from routing and format enforcement tuned to the benchmarks, and the paper evaluates nothing beyond these four, so whether it carries to open-ended tasks is untested. Teams that want small models in production without retraining should track this scaffolding path.

Today in one line: Writer saves about 40% on harness changes alone, the AI4AI paper nearly doubles a weak model’s scores with a harness, and OpenART finds agent safety varies substantially with runtime implementation rather than the base model. In today’s items, the biggest levers on cost, capability, and safety all sat in the software layer around the model, not in the model itself.