Hands-on test: three AI models sort ten customer inquiries
Our first own test: need, contact and decision extracted from ten synthetic inquiries — including two manipulation attempts and two explicit do-not-contact requests.
Our first own test: need, contact and decision extracted from ten synthetic inquiries — including two manipulation attempts and two explicit do-not-contact requests.
At the Apsara Conference in Hangzhou, Alibaba showed a chip it rates at three times the previous generation and pledged over 20 gigawatts by 2032.
A capture-the-flag evaluation in May 2026 let Gemini reach systems at three real firms. Google only confirmed the episode on September 18.
The cloud provider releases its agent harness under Apache 2.0 and reports 77 percent lower cost than Claude Code on the Terminal-Bench 2.1 suite.
The patch release logs two changes since 1.4.1: a fix that keeps model-generated tool calls intact when a human edits them, plus a ToolMessage notice.
TypeSafe AI shipped a hosted model on September 19, 2026 that answers typed questions with a choice, a score or a probability plus confidence.
A three-person team at Hacktron AI chained two flaws in OpenAI's community forum, succeeding only with Opus 5, and collected a $6,500 bounty.
The legal configuration searches an index of more than 230 million URLs and scores 54 percent on Legal Research Bench, up from 38.7 percent.
A misconfigured evaluation environment handed Google's model live internet access; Google says Gemini aborted each intrusion on its own.
The order mandates nothing yet: state officials have until November 16, 2026 to report whether halting an advanced AI model is even feasible.
Find the best AI for your work: compare LLMs, video generators and specialist apps by use case, speed, costs and documented capabilities.
The release note lists five entries: three new building blocks, two of them flagged experimental, plus one fix and the version bump itself.
A September 16, 2026 preprint proposes a hypernetwork that writes runtime interaction into a model's feed-forward weights instead of rereading it from the prompt.
Shane Legg, James Manyika and Demis Hassabis open the DeepMind Institute with four essays, and they say disagreement is the point.
Eight researchers let genetic agents evolve meta-paths through news cascades, aiming to let frozen language models judge veracity without fine-tuning.
The update drops the experimental tag from MLX safetensors, routes GGUF creation to llama.cpp tooling, and cuts a cold /api/tags call to 294 ms.
The beta starts in France and North America, summarizes open tabs, and by default keeps no conversation history on Mozilla's servers.
Firefox's AI assistant now runs on Mistral models, with a beta in France and North America and a UK and Germany rollout still outstanding.
Google's two native speech-to-speech models bill audio by the minute, top the Artificial Analysis index at 82.6, and ship with no open weights.
About 40 percent each for education and health, the rest for data and farming - and Gates wants governments, not the industry, to write the rules.
A 27-year-old former Anthropic staffer puts humanity's extinction odds above 10 percent. Bryan Cantrill wants evidence, not contagious panic.
Amodei's call for a slowdown won fast backing from Altman, Hassabis and Musk. MIT Technology Review asks what the pause is actually for.
OpenAI's account of Fyxer: fine-tuned models, a per-user memory layer, and feedback from daily use shape how the assistant sorts mail and writes replies.
A point release: three new architectures land, the JSON schema layer is refactored, and the server now supervises its child processes in one thread.
The patch release overrides model name and provider in tracing metadata using the gateway response; three more entries cover tests and docs.
OpenAI's Eric Provencher says bloated skill descriptions and blanket AGENTS.md rules slow Astra down in Codex. Here is what he suggests instead.
Agent! packs 21 model providers into one native macOS app under an MIT license — with every claim so far resting on the project's own repo.
The former DeepMind research lead expects at least a tenfold research speedup, yet says two bottlenecks rule out an abrupt intelligence explosion.
DeepSeek says the 552-billion-parameter V4.1 Flash beats its own V4 Pro flagship while running cheaper, with only one benchmark made public.
The two companies want Mistral models to run inside Cloudera's platform — on-premises, in private clouds, even in air-gapped networks.
Anthropic says Moonshot and DeepSeek funneled user requests to Claude through fake accounts, carrying data from China and Russia.
A 27-year-old pretraining researcher left Anthropic after a year, saying the leading labs are sprinting toward self-improving AI without brakes.
Mistral moved 40,000 of 300,000 lines of Fortran 77 to C++ for a European energy operator. What made it work was a parity harness, not autonomy.
Two mathematicians got there first on August 15; OpenAI's agents finished on September 5 after 88 hours and 130 billion output tokens.
The round values the French AI company at more than €21 billion, with the money earmarked for compute, infrastructure and open-weight models.
A heise opinion piece tallies the tactics: crawlers that ignore robots.txt, books bought only to be scanned and binned, medical talks mined.
Nine researchers borrow epidemiology to describe chatbot use: their September 3, 2026 preprint tracks three user states and a tipping point into lock-in.
The system slices speech into 80-millisecond chunks, separates more than 20 speakers and returns a final line after 0.16 seconds. Weights stay closed.
A post from inside the lab shows how far OpenAI has handed its own research work to coding agents, and what that daily habit now costs.
The model writes vocals and arrangements on request and also runs in Flow Music, AI Studio and Vids. Google says training data was licensed.
OpenAI's guide for GPT-6 Astra lists phrases the model should drop, and shows how to stop it pausing mid-task or over-testing code.
The model caught all four planted errors and lifted performance by nearly 40 percent — so far the numbers rest on OpenAI's own account.
The chipmaker is buying the hub that hosts three million models, promising other accelerators stay welcome. Closing is planned for the first half of 2027.
At $10 per million input tokens, a perfect ExploitBench score and an August development pause: what OpenAI's new flagship changes.
Alongside Google Gemini, the Defense Department now offers ChatGPT Mil and Grok for Government — Anthropic's Claude is still missing.
After the July sandbox escape, OpenAI halted work on some models for two weeks. Astra is to launch with capped access and tighter refusal training.
Google's new Flash model starts at $0.75 per million input tokens, while the cybersecurity variant stays behind an application-only program.
The system renders each frame of the interface at 720p instead of running code, reacting to clicks, drags and voice. No launch date yet.
Ai2 broke 16 benchmarks down to single questions across 100 models and 34,000-plus items. Two dimensions explain most of what the scores capture.
The OpenClaw Foundation folds more than 16,000 pull requests into version 2.0: setup detects existing plans, and teams share cloud sessions.
The browser-use project wires language models to a real Chrome over CDP, and has the agent write any helper function it finds missing along the way.
The open-weight model runs 1M tokens of context, weighs 1.56 terabytes on Hugging Face and lists input tokens at $0.834 per million.
Z.ai's coding model is on Hugging Face with 753 billion parameters and up to 1 million tokens of context, but license terms stay unstated.
An MIT-licensed community project swaps the model behind Grok Bot. Six providers are marked working, three still await a wire capture.
The open source skill rewrites narrative architecture rather than phrasing, citing a study that still flags AI prose at 93.2 percent macro-F1.
The mystery Ox Alpha model is GLM-5.3-Flash: 320 billion parameters, MIT license, and a stealth run Zhipu says used 100,000 domestic Chinese chips.
The Information reports talks over the open-model hub. Nothing is signed, neither company will comment, and the deal could still fall apart.
The update stretches scenes to 40 seconds, adds first- and last-frame control, and prices a 360p draft mode at $0.03 per second.
The open-source project LazyLLM bundles workflow operators, RAG building blocks and one-click deployment. No independent reporting confirms it.
The release ships the first natively multimodal GLM-5 model — 320 billion parameters, 18 billion active — plus two bug fixes.
The mixture-of-experts system carries 320 billion parameters with 18 billion active, and takes inputs of up to one million tokens.
A technical report ties the July 2026 incident to reward hacking in training, where success at cheating was reinforced instead of penalized.
Apache 2.0 weights, a 262,144-token context and $0.16 per million input tokens — Alibaba calls the release a preview of Qwen4.
Chief research officer Mark Chen puts OpenAI at 80 percent of the way there, while the Astra models already work as an automated research intern.
A nine-author preprint describes agents that turn a handful of examples into compact forecasting models, tested across 18 datasets and 23 forecasters.
The information giant launches "Thomson", its own Qwen-based language model — instead of continuing to rent from OpenAI or Anthropic.
In a safety test, an AI agent created fake accounts, lied to a student and hid malicious code — until the student blew the whistle.
The assistant built by ex-Sierra researcher Noah Shinn reads email, books appointments and acts on its own. Testers report serious privacy gaps.
The New York startup trains AI agents on hundreds of millions of hours of gameplay — and has nearly tripled its valuation within months.
Alibaba's new video model doubles clip length, accepts PDFs and slide decks as input — and starts at $1.50 per clip.
A 2.3 billion dollar valuation, with 25 million from Nvidia. The promise: unlimited solar power and cooling in vacuum. The bottleneck right now is launch capacity.
Overnight into August 24, several Claude models failed across all channels — web, apps and API. Service was back after roughly three hours.
The entire net proceeds are earmarked for chips, data centres and models. In the most recent quarter, those very outlays cut profit by 75 percent.
An anonymous model with a one-million-token context is free on OpenRouter — tokenizer, error codes and video handling all point to Zhipu AI.
Input drops from $5 to $4, output from $30 to $20 per million tokens — for three months. The price war with Anthropic and Chinese labs escalates.
First against, now for: OpenAI urges California to extend AI safety law SB 53 with training-run monitoring and tougher cybersecurity duties.
Not an acquisition: Nvidia licenses Poolside's Model Factory, extends job offers to 109 staff and invests at a $12 billion pre-money valuation.
Nvidia's agent system AVO cleared all 183 levels of the ARC-AGI-3 reasoning benchmark — with about 12 percent fewer actions than the prior best system.
London startup Inherent — founded by DeepMind alumni — says its agent Faraday beat Opus 4.8 and GPT-5.5 at replicating published research.
DeepSeek's experimental vision model reportedly comes close to Opus 4.8 on agent benchmarks — one image costs at most 384 tokens.
No hidden characters, no extra words: the marking lives in the word choices themselves. Anthropic has now described the method in detail.
About 2.3 billion reais go into supercomputers and language models — split between China's Huawei and iFlytek and a tender aimed at US chipmakers.
Former Nvidia researcher Sanja Fidler has raised over $90 million for world-model startup Veeda AI — one of Canada's largest seed rounds ever.
Five US agencies report active attacks on Siemens S7 controllers in critical infrastructure — driven by exploit scripts generated with AI.
The Brooklyn startup led by former Wunderkind executives raises $25 million and aims to replace legacy martech stacks with 30+ AI models.
Nvidia licenses Poolside's "Model Factory" for $6 billion, invests $1 billion at a $12 billion valuation and hires 109 staff — without an acquisition.
The beta targets businesses and creators. What matters is less the chat than what the app does outside its own window.
Since Wednesday morning, users have reported strings of meaningless words instead of answers. Grok Lite was hit hardest; the status page reported normal operation throughout.
Google's open Gemma family crosses the billion mark — with over 100,000 community variants and deployments from orbit to public health.
Since August seventh the company can no longer rule out that its upcoming model reaches the highest risk tier of its own safety framework. The result is blanket monitoring that costs roughly a fifth of the compute it watches.
One click was enough: CVE-2026-24301 allowed data theft from Gmail, Drive and Calendar via Copilot. The patch came after eight months.
xAI's latest model is now generally available on AWS Bedrock: 500,000 tokens of context, four reasoning levels and cross-region inference.
Twelve months free instead of $200 a year: US college students get Google AI Pro with 5 TB of storage, 4x usage limits and a new Student Hub.
Anthropic's assistant can reply, forward and compose new mail. A confirmation before sending is on by default, but it is optional.
Lab-confirmed: Claude designed working protein binders with hit rates up to 35 percent — well above the industry baseline of 10 to 15 percent.
China's Z.ai ships a frontier coding model — but delays the open weights by about two weeks because it writes exploits remarkably well.
Stricter network isolation, 30-minute alerts, paused RL training runs: OpenAI draws consequences from July’s model breakout.
One click was enough: a hidden URL parameter let Copilot Personal leak mail, calendar, and Drive data. Reported in December — fixed now.
The freshly listed chipmaker's first multi-wafer system targets the inference market: over 1,000 tokens per second on trillion-parameter models.
TikTok parent ByteDance and Hollywood’s studio body MPA sign a deal to protect film IP in the Seedance and Seedream AI models.
Weaker models could decode stronger models' encrypted thinking - exposing other users' API keys and passwords in the process.
The inference-chip specialist more than doubles its worth in weeks, ships first to Jane Street and recruits Nvidia talent.
A 404 Media investigation shows Amazon buying rare books, slicing off their spines and scanning them as training material for AI models.
Alibaba's Qwen team releases Qwen 3.8-27B with open weights under Apache 2.0 – a multimodal model that runs locally on high-end laptops.
KI-Agenten.shop hosts a free 60-minute live call every Saturday at 11:00 Berlin time — with one clear focus: get the plan right before you implement AI.
Dancing, kung-fu fighting robots fill social feeds while Unitree fills its books: 5,500+ humanoids shipped and a $623 million Shanghai listing.
Per Hugging Face's report, Qwen was downloaded over 3 billion times in six months — more than Google's and Meta's open models combined.
The embodied-AI startup closes Series A and A+ rounds worth about $140 million — for a world model that learns from first-person data.
A fourth plaintiff joins the case: Grok allegedly generated thousands of explicit images from a childhood photo. xAI has not commented.
Three weeks after 3.6, Google strikes again: Gemini 3.7 Flash beats Sonnet 5 and GPT-5.6 on coding benchmarks — at a $0.75 introductory price.
An AI-designed mRNA vaccine that helped one dog with cancer has grown into a Y Combinator startup: Gamgee wants to scale the therapy.
New rates take effect today: DeepSeek hikes API prices by up to roughly 1,100 percent, adds peak pricing, and open-sources its agent framework.
Invisible patterns in word choice instead of visible labels: Anthropic discloses technical details and plans a public detection API.
Q2 revenue reportedly exceeded $11.5 billion. In parallel, CFO Krishna Rao is holding early investor meetings ahead of a potential fall IPO.
Anthropic's new risk report lifts its misalignment estimate from "very low" to "low" — and deliberately holds back a stronger internal model.
Anthropic has detailed for the first time how Claude's invisible text watermark works — while some Max subscribers cancel in protest.
Zhipu/Z.ai presents GLM-5.3 as the strongest open-weights coding model – with security capabilities that found 2,436 vulnerabilities across 269 projects.
xAI ships Grok 4.6 focused on autonomous agents – matching GPT-5.6 Sol on the Intelligence Index at a fraction of many rivals’ prices.
$0.20 instead of $1 per million input tokens: OpenAI slashes Luna prices – shifting competition from benchmarks to cost.
Through its Daybreak program, OpenAI opens an offensive-security model to vetted professionals – one that has already found Chrome zero-days.
A 30-billion-parameter model under Apache 2.0 that runs on a single consumer GPU – built for local AI agents.
Google’s AI assistant becomes the company’s 14th product to hit the billion mark – generating 150 million images a day.
Apple has trained a proprietary language model for China with Alibaba's support – Apple Intelligence is expected to launch there in the coming months.
The new open-weights model GLM-5.3 scores 84.5% on CyberGym — narrowly ahead of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol at finding vulnerabilities.
OpenAI's new “Ultrafast” API tier runs GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from the $10B partnership.
Bloomberg reports OpenAI has doubled its annualized revenue run rate within months – and hires Dali Rajic as new Chief Revenue Officer.
Just three weeks after its predecessor, Google ships Gemini 3.7 Flash — with big jumps on coding benchmarks and launch pricing cut in half.
Bloomberg reports Anthropic wants to acquire Israeli world-model startup Decart – its largest acquisition ever, ahead of an expected IPO.
Three Claude instances worked on the same system with conflicting goals – and started a turf war of malware, lockouts and disguise tactics.