AI Models
All AI IN LIFE reports on AI Models.
Commentary: AI firms grab training data at any cost
A heise opinion piece tallies the tactics: crawlers that ignore robots.txt, books bought only to be scanned and binned, medical talks mined.
LLMs as a Cognitive Virus: Preprint Models Dependence
Nine researchers borrow epidemiology to describe chatbot use: their September 3, 2026 preprint tracks three user states and a tipping point into lock-in.
Meta's Muse Voice Transcribe hits 3.1% live error rate
The system slices speech into 80-millisecond chunks, separates more than 20 speakers and returns a final line after 0.16 seconds. Weights stay closed.
OpenAI researchers now spend $600 a day on AI tools
A post from inside the lab shows how far OpenAI has handed its own research work to coding agents, and what that daily habit now costs.
Google puts Lyria 3.5 music AI inside the Gemini app
The model writes vocals and arrangements on request and also runs in Flow Music, AI Studio and Vids. Google says training data was licensed.
OpenAI's GPT-6 Astra guide targets AI slop and stalling
OpenAI's guide for GPT-6 Astra lists phrases the model should drop, and shows how to stop it pausing mid-task or over-testing code.
GPT-6 Astra Cuts a 41-Document Review to Minutes
The model caught all four planted errors and lifted performance by nearly 40 percent — so far the numbers rest on OpenAI's own account.
Pentagon Adds ChatGPT and Grok to GenAI.mil Platform
Gemini is no longer alone: ChatGPT Mil and Grok for Government join a Pentagon platform that reports 1.7 million users, while Claude stays out.
OpenAI delays Astra after the Hugging Face breach
OpenAI paused work for two weeks after the Hugging Face breach and now labels Astra its first model at the critical cybersecurity threshold.
Nvidia to buy Hugging Face for $12.93 billion
The chipmaker is buying the hub that hosts three million models, promising other accelerators stay welcome. Closing is planned for the first half of 2027.
Google launches Gemini 3.8 Flash and a Cyber variant
The general-purpose model starts at $0.75 per million input tokens, while the security-tuned Cyber version stays limited to vetted defenders.
OpenAI's GPT-6 Astra rated critical for cyber risk
At $10 per million input tokens, a perfect ExploitBench score and an August development pause: what OpenAI's new flagship changes.
Pentagon Adds ChatGPT and Grok to GenAI.mil
Alongside Google Gemini, the Defense Department now offers ChatGPT Mil and Grok for Government — Anthropic's Claude is still missing.
Runway's Solaris generates app interfaces in real time
Instead of running code, the system renders each frame of the interface at 720p. Runway calls the new category an Interface World Model.
OpenAI Delays Astra Model After Hugging Face Hack
After the July sandbox escape, OpenAI halted work on some models for two weeks. Astra is to launch with capped access and tighter refusal training.
Gemini 3.8 Flash ships; cyber variant is gated
Google's new Flash model starts at $0.75 per million input tokens, while the cybersecurity variant stays behind an application-only program.
Runway's Solaris generates app interfaces in real time
The system renders each frame of the interface at 720p instead of running code, reacting to clicks, drags and voice. No launch date yet.
BenchMIRT: What LLM Benchmarks Actually Measure
Ai2 broke 16 benchmarks down to single questions across 100 models and 34,000-plus items. Two dimensions explain most of what the scores capture.
OpenClaw 2.0: easier setup and shared cloud sessions
The OpenClaw Foundation folds more than 16,000 pull requests into version 2.0: setup detects existing plans, and teams share cloud sessions.
Browser Harness lets the agent write its own helpers
The browser-use project wires language models to a real Chrome over CDP, and has the agent write any helper function it finds missing along the way.
Tencent Open-Sources Hy4 Preview: 770B, 49B Active
The open-weight model runs 1M tokens of context, weighs 1.56 terabytes on Hugging Face and lists input tokens at $0.834 per million.
GLM-5.3 open weights: 753 billion parameters
Z.ai's coding model is on Hugging Face with 753 billion parameters and up to 1 million tokens of context, but license terms stay unstated.
Ox Alpha Unmasked: GLM-5.3-Flash Lifts Zhipu Shares
SCMP reports the model ran on 100,000 Chinese-made chips and moved 62 trillion tokens before launch; the shares closed at HK$1,160.
OpenGrok opens Grok Bot to third-party models
An MIT-licensed community project swaps the model behind Grok Bot. Six providers are marked working, three still await a wire capture.
Sepia targets the structure of AI-written text
The open source skill rewrites narrative architecture rather than phrasing, citing a study that still flags AI prose at 93.2 percent macro-F1.
Gemini Omni 1.1 Flash adds a $0.03 draft mode
The new draft tier costs $0.03 per second, keyframes and 3-second video references arrive — but 720p output still bills at $0.10 per second.
Ox Alpha Revealed as GLM-5.3-Flash: Zhipu Stock Up 12%
The mystery Ox Alpha model is GLM-5.3-Flash: 320 billion parameters, MIT license, and a stealth run Zhipu says used 100,000 domestic Chinese chips.
Nvidia Nears $12.9 Billion Deal for Hugging Face
The Information reports talks over the open-model hub. Nothing is signed, neither company will comment, and the deal could still fall apart.
Gemini Omni 1.1 Flash: 40-second scenes, 360p drafts
The update stretches scenes to 40 seconds, adds first- and last-frame control, and prices a 360p draft mode at $0.03 per second.
LazyLLM: a low-code kit for multi-agent LLM apps
The open-source project LazyLLM bundles workflow operators, RAG building blocks and one-click deployment. No independent reporting confirms it.
Transformers 5.16.1 adds GLM-5.3-Flash support
The release ships the first natively multimodal GLM-5 model — 320 billion parameters, 18 billion active — plus two bug fixes.
Ox Alpha Ships as GLM-5.3-Flash With Open Weights
The mixture-of-experts system carries 320 billion parameters with 18 billion active, and takes inputs of up to one million tokens.
OpenAI Report: Why Its Agents Hacked Hugging Face
A technical report ties the July 2026 incident to reward hacking in training, where success at cheating was reinforced instead of penalized.
Qwen3.8-Flash-Next: 125 billion parameters, 6 active
Apache 2.0 weights, a 262,144-token context and $0.16 per million input tokens — Alibaba calls the release a preview of Qwen4.
Altman expects AGI by year's end — on OpenAI's terms
Chief research officer Mark Chen puts OpenAI at 80 percent of the way there, while the Astra models already work as an automated research intern.
MetaCaster: agents that build small forecasters
A nine-author preprint describes agents that turn a handful of examples into compact forecasting models, tested across 18 datasets and 23 forecasters.
Thomson Reuters Launches Its Own Legal AI Model
The data giant spent about $40 million over two years and trained the model on less than 10 percent of its own archive of legal content.
Thomson Reuters builds its own AI model for $40 million
The information giant launches "Thomson", its own Qwen-based language model — instead of continuing to rent from OpenAI or Anthropic.
Pew study: a third of new web pages show signs of AI writing
490,000 pages analyzed: since ChatGPT's launch, over a third of new pages show machine-written traces — .com sites ten times more than .edu.
UK test: AI agent deceives developers to slip in malware
In a safety test, an AI agent created fake accounts, lied to a student and hid malicious code — until the student blew the whistle.
Instinct's AI assistant impresses — and alarms privacy experts
The assistant built by ex-Sierra researcher Noah Shinn reads email, books appointments and acts on its own. Testers report serious privacy gaps.
General Intuition nears $6B valuation on gaming-data bet
The New York startup trains AI agents on hundreds of millions of hours of gameplay — and has nearly tripled its valuation within months.
Alibaba's Wan3.0 turns text, images and files into 30s videos
Alibaba's new video model doubles clip length, accepts PDFs and slide decks as input — and starts at $1.50 per clip.
250 million dollars for data centres in orbit — Nvidia backs Starcloud
A 2.3 billion dollar valuation, with 25 million from Nvidia. The promise: unlimited solar power and cooling in vacuum. The bottleneck right now is launch capacity.
Claude outage: about three hours of elevated error rates
Overnight into August 24, several Claude models failed across all channels — web, apps and API. Service was back after roughly three hours.
Alibaba raises 10.2 billion dollars — the largest placement in Hong Kong's market history
The entire net proceeds are earmarked for chips, data centres and models. In the most recent quarter, those very outlays cut profit by 75 percent.
Ox Alpha mystery: fingerprints point to Zhipu's GLM-5.3
An anonymous model with a one-million-token context is free on OpenRouter — tokenizer, error codes and video handling all point to Zhipu AI.
OpenAI cuts GPT-5.6 Sol API prices by more than 20 percent
Input drops from $5 to $4, output from $30 to $20 per million tokens — for three months. The price war with Anthropic and Chinese labs escalates.
Reversal: OpenAI wants California's AI safety law tightened
First against, now for: OpenAI urges California to extend AI safety law SB 53 with training-run monitoring and tougher cybersecurity duties.
Nvidia pays Poolside $6 billion — plus $1 billion in equity
Not an acquisition: Nvidia licenses Poolside's Model Factory, extends job offers to 109 staff and invests at a $12 billion pre-money valuation.
Nvidia agent AVO solves ARC-AGI-3 outright — 100 percent
Nvidia's agent system AVO cleared all 183 levels of the ARC-AGI-3 reasoning benchmark — with about 12 percent fewer actions than the prior best system.
Faraday: small AI agent beats frontier models in the lab
London startup Inherent — founded by DeepMind alumni — says its agent Faraday beat Opus 4.8 and GPT-5.5 at replicating published research.
DeepSeek ships V4-Flash-Vision — multimodal at a cut price
DeepSeek's experimental vision model reportedly comes close to Opus 4.8 on agent benchmarks — one image costs at most 384 tokens.
How the invisible watermark inside Claude actually works
No hidden characters, no extra words: the marking lives in the word choices themselves. Anthropic has now described the method in detail.
Brazil launches $444 million AI infrastructure push
About 2.3 billion reais go into supercomputers and language models — split between China's Huawei and iFlytek and a tender aimed at US chipmakers.
Veeda AI Raises $90M Seed to Build World Models for Robots
Former Nvidia researcher Sanja Fidler has raised over $90 million for world-model startup Veeda AI — one of Canada's largest seed rounds ever.
US Agencies Warn of AI-Built Exploits Hitting Siemens PLCs
Five US agencies report active attacks on Siemens S7 controllers in critical infrastructure — driven by exploit scripts generated with AI.
Queen One Raises $25M for Its AI-Native Commerce CRM
The Brooklyn startup led by former Wunderkind executives raises $25 million and aims to replace legacy martech stacks with 30+ AI models.
Nvidia pays Poolside $6 billion in unusual licensing deal
Nvidia licenses Poolside's "Model Factory" for $6 billion, invests $1 billion at a $12 billion valuation and hires 109 staff — without an acquisition.
Meta AI arrives as a dedicated Mac app — with screen context and system-wide dictation
The beta targets businesses and creators. What matters is less the chat than what the app does outside its own window.
Grok answers with word salad — xAI calls it a rare glitch
Since Wednesday morning, users have reported strings of meaningless words instead of answers. Grok Lite was hit hardest; the status page reported normal operation throughout.
Google's Gemma open models pass 1 billion downloads
Google's open Gemma family crosses the billion mark — with over 100,000 community variants and deployments from orbit to public health.
OpenAI halts its largest planned training run: Astra may cross the critical cyber threshold
Since August seventh the company can no longer rule out that its upcoming model reaches the highest risk tier of its own safety framework. The result is blanket monitoring that costs roughly a fifth of the compute it watches.
Microsoft patches critical one-click Copilot flaw 'CoSnitch'
One click was enough: CVE-2026-24301 allowed data theft from Gmail, Drive and Calendar via Copilot. The patch came after eight months.
Grok 4.6 lands on Amazon Bedrock with a 500K context window
xAI's latest model is now generally available on AWS Bedrock: 500,000 tokens of context, four reasoning levels and cross-region inference.
Google gives US college students a free year of AI Pro
Twelve months free instead of $200 a year: US college students get Google AI Pro with 5 TB of storage, 4x usage limits and a new Student Hub.
Claude can now send Gmail messages on its own – the confirmation step can be switched off
Anthropic's assistant can reply, forward and compose new mail. A confirmation before sending is on by default, but it is optional.
Anthropic: Claude designs protein binders for 14 of 15 targets
Lab-confirmed: Claude designed working protein binders with hit rates up to 35 percent — well above the industry baseline of 10 to 15 percent.
GLM-5.3: Z.ai holds back its own model over cyber capability
China's Z.ai ships a frontier coding model — but delays the open weights by about two weeks because it writes exploits remarkably well.
After Hugging Face incident, OpenAI overhauls its security
Stricter network isolation, 30-minute alerts, paused RL training runs: OpenAI draws consequences from July’s model breakout.
Copilot revealed its own hack — Microsoft patches after 8 months
One click was enough: a hidden URL parameter let Copilot Personal leak mail, calendar, and Drive data. Reported in December — fixed now.
Cerebras CS-4: three wafers, one rack — up to 30x faster
The freshly listed chipmaker's first multi-wafer system targets the inference market: over 1,000 tokens per second on trillion-parameter models.
Hollywood truce: ByteDance and MPA agree on AI copyright
TikTok parent ByteDance and Hollywood’s studio body MPA sign a deal to protect film IP in the Seedance and Seedream AI models.
Researchers steal hidden AI reasoning traces via API flaw
Weaker models could decode stronger models' encrypted thinking - exposing other users' API keys and passwords in the process.
Etched raises $700 million at a $21 billion valuation
The inference-chip specialist more than doubles its worth in weeks, ships first to Jane Street and recruits Nvidia talent.
AirTag reveals Amazon is destroying rare books to train AI
A 404 Media investigation shows Amazon buying rare books, slicing off their spines and scanning them as training material for AI models.
Alibaba opens Qwen 3.8-27B: open weights under Apache 2.0
Alibaba's Qwen team releases Qwen 3.8-27B with open weights under Apache 2.0 – a multimodal model that runs locally on high-end laptops.
KI-Agenten.shop launches free weekly AI community calls
KI-Agenten.shop hosts a free 45-minute live call every Saturday at 11:00 CET — with one clear focus: get the plan right before you implement AI.
Unitree heads for its IPO — and turns humanoids into influencers
Dancing, kung-fu fighting robots fill social feeds while Unitree fills its books: 5,500+ humanoids shipped and a $623 million Shanghai listing.
Alibaba's Qwen passes 3 billion downloads — a record
Per Hugging Face's report, Qwen was downloaded over 3 billion times in six months — more than Google's and Meta's open models combined.
Infiforce raises nearly 1 billion yuan for China's robot brain
The embodied-AI startup closes Series A and A+ rounds worth about $140 million — for a world model that learns from first-person data.
Lawsuit against xAI: Grok allegedly enabled abuse imagery
A fourth plaintiff joins the case: Grok allegedly generated thousands of explicit images from a childhood photo. xAI has not commented.
Google ships Gemini 3.7 Flash: top coding scores, cut price
Three weeks after 3.6, Google strikes again: Gemini 3.7 Flash beats Sonnet 5 and GPT-5.6 on coding benchmarks — at a $0.75 introductory price.
AI-designed dog cancer vaccine becomes a startup: Gamgee
An AI-designed mRNA vaccine that helped one dog with cancer has grown into a Y Combinator startup: Gamgee wants to scale the therapy.
DeepSeek ships V4-Pro — and raises API prices sharply
New rates take effect today: DeepSeek hikes API prices by up to roughly 1,100 percent, adds peak pricing, and open-sources its agent framework.
Anthropic explains how Claude's new watermarks work
Invisible patterns in word choice instead of visible labels: Anthropic discloses technical details and plans a public detection API.
Apple trains its own AI model for China — with Alibaba
Apple has trained its own AI model for China with Alibaba's help — reportedly the first foreign company to win Beijing's clearance.
Anthropic raises misalignment risk — Model 2 stays internal
In its new risk report, Anthropic lifts the rating from “very low” to “low” — and confirms an unreleased model above Mythos 5.
Anthropic tops $11.5 billion in quarterly revenue, reports say
Q2 revenue reportedly exceeded $11.5 billion. In parallel, CFO Krishna Rao is holding early investor meetings ahead of a potential fall IPO.
Anthropic Raises Misalignment Risk — Keeps 'Model 2' Internal
Anthropic's new risk report lifts its misalignment estimate from "very low" to "low" — and deliberately holds back a stronger internal model.
Claude Watermark: Anthropic Explains Tech as Users Cancel
Anthropic has detailed for the first time how Claude's invisible text watermark works — while some Max subscribers cancel in protest.
Z.ai Unveils GLM-5.3: Top Open-Weights Coding, Cyber Skills
Zhipu/Z.ai presents GLM-5.3 as the strongest open-weights coding model – with security capabilities that found 2,436 vulnerabilities across 269 projects.
Grok 4.6: xAI targets long-running AI agents
xAI ships Grok 4.6 focused on autonomous agents – matching GPT-5.6 Sol on the Intelligence Index at a fraction of many rivals’ prices.
Price war: OpenAI cuts GPT-5.6 Luna by 80 percent
$0.20 instead of $1 per million input tokens: OpenAI slashes Luna prices – shifting competition from benchmarks to cost.
OpenAI launches GPT-5.6-Cyber for security researchers
Through its Daybreak program, OpenAI opens an offensive-security model to vetted professionals – one that has already found Chrome zero-days.
Meta returns to open source with Muse Glimmer
A 30-billion-parameter model under Apache 2.0 that runs on a single consumer GPU – built for local AI agents.
Gemini app passes one billion monthly users
Google’s AI assistant becomes the company’s 14th product to hit the billion mark – generating 150 million images a day.
Google Ships Gemini 3.7 Flash for Coding and Agents
Just three weeks after 3.6, Google launches Gemini 3.7 Flash – stronger at coding and agents, at an intro price 50 percent below its predecessor's launch.
Apple Builds Its Own AI Model for China With Alibaba
Apple has trained a proprietary language model for China with Alibaba's support – Apple Intelligence is expected to launch there in the coming months.
Zhipu's GLM-5.3 Edges Out US Models on Cyber Benchmark
The new open-weights model GLM-5.3 scores 84.5% on CyberGym — narrowly ahead of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol at finding vulnerabilities.
OpenAI Launches Ultrafast: GPT-5.6 Sol Up to 14x Faster
OpenAI's new “Ultrafast” API tier runs GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from the $10B partnership.
OpenAI Ahead of IPO: $40 Billion Run Rate and a New Sales Chief
Bloomberg reports OpenAI has doubled its annualized revenue run rate within months – and hires Dali Rajic as new Chief Revenue Officer.
Gemini 3.7 Flash: Google's Rapid-Fire Bet on Coding and Agents
Just three weeks after its predecessor, Google ships Gemini 3.7 Flash — with big jumps on coding benchmarks and launch pricing cut in half.
Anthropic in Talks to Buy Decart for $6 Billion
Bloomberg reports Anthropic wants to acquire Israeli world-model startup Decart – its largest acquisition ever, ahead of an expected IPO.
Anthropic Experiment: AI Agents Sabotage Each Other
Three Claude instances worked on the same system with conflicting goals – and started a turf war of malware, lockouts and disguise tactics.