Anthropic Details Claude Content Watermark Update
Anthropic is reported to have shared further details on how content generated with Claude is watermarked.
- Anthropic
- Claude
- AIWatermark
- ContentProvenance
- AIPolicy
New AI model releases and research
130 articles
Anthropic is reported to have shared further details on how content generated with Claude is watermarked.
Chinese AI firm Zhipu claims its new model outperforms Anthropic and OpenAI at finding bugs in code.
Anthropic reportedly issued guidance urging developers to avoid unnecessary token consumption when using its Claude Code AI coding tool.
Alibaba's Qwen has reportedly moved ahead of Meta's Llama and Google in the race for leadership among open-weight AI models.
A HackerNoon piece walks through patterns for smoothly streaming LLM responses into React Native chat UIs instead of hand-rolling them.
Yahoo Tech argues Anthropic's new model tops a rival on benchmarks, yet everyday users are unlikely to feel the difference.
OpenAI's GPT-5.6 Sol reportedly runs up to 14× faster in a new 'Ultrafast' mode, targeting latency-sensitive real-world deployments.
OpenAI and Anthropic are reportedly slashing prices to counter fast-rising, low-cost Chinese AI models competing on value.
Publishing veteran Tim O'Reilly says he backs AI even as it erodes his own business — but only when it's open source.
Global communications firm Ruder Finn has unveiled a new website powered by a large language model, signaling the marketing and PR industry's move to put generative AI at the customer-facing front line.
xAI shipped Grok 4.6 on August 7, 2026, keeping its 1.5-trillion-parameter V9 foundation while boosting performance through improved supervised fine-tuning and reinforcement learning — enough to surpass Claude Opus 4.8 on the SWE-Marathon benchmark.
Researchers unveiled a technique to extract reasoning traces from Claude, GPT, and Gemini. They say the findings suggest some Chinese AI models may have been trained on leading US systems.
Alibaba released Qwen Image 3.0 Pro on August 5, its third-generation image-generation model targeting commercial and document use cases — but with no benchmarks, no model card, and no public API, in a notable transparency retreat.
Meta launched Muse Glimmer on August 10, a 30B-parameter open-weight agentic AI model under Apache 2.0 that runs on a single consumer GPU — its first purpose-built open agent model.
Nikkei Asia reports that Moonshot AI's founder turned down a job offer from Apple and instead built a Chinese AI startup positioned as a rival to Anthropic.
A commentary argues that Meta's back-and-forth on releasing open-weight models is unsettling observers of the open AI ecosystem.
Model distillation — transferring a large model's knowledge into a smaller one — is reportedly becoming a new focal point in the US-China AI race.
Anthropic has reportedly changed its Claude Code agent to run more autonomously by default, reducing step-by-step approvals so it can carry out coding tasks on its own.
Wired highlights Meetily, a free and open-source tool that transcribes and summarizes meetings without a subscription.
An OpenAI executive reportedly predicted that the company's Codex coding tool could become obsolete within two to three months — a sign of how fast AI coding tools are turning over.
Anthropic has reportedly switched its Claude Code coding tool to Auto Mode by default, aiming to reduce developer mistakes caused by careless approvals.
DeepSeek released V4-Flash-0731 on July 31, 2026 — same 284B/13B MoE architecture as Flash-Preview, but with post-training rebuilt from scratch. The result: Flash pricing with benchmark scores that beat DeepSeek's own flagship V4-Pro-Preview across all nine agent and coding tasks.
NVIDIA released NOOA, an object-oriented Python framework that reduces an AI agent to a single class definition — and claims 82.2% on SWE-bench Verified with GPT-5.5 using half the tokens of prior state-of-the-art.
TikTok parent ByteDance is reportedly training a massive new AI model aimed at rivaling Anthropic's 'Mythos'.
TikTok owner ByteDance is aiming to build a large AI model approaching the capability of Anthropic's top system, according to a Financial Times report cited by Reuters.
According to a Financial Times report, ByteDance is developing a large-scale AI model aimed at approaching the performance of Anthropic's top-tier 'Mythos' model.
Demis Hassabis steps back from the CEO role at Google DeepMind to become chairman and Alphabet chief scientist, focusing on AGI strategy. CTO Koray Kavukcuoglu is promoted to SVP to run day-to-day operations and Gemini development.
xAI launched Grok Voice Think Fast 2.0, cutting first-audio response time to 0.70 seconds and reducing inference token usage by 60%. The model became the new default voice for Grok on August 5.
Alibaba released Qwen3.8-Max on August 3, a 2.4 trillion-parameter mixture-of-experts model that topped the Chinese leaderboard on Arena.AI and ranked second globally on vision tasks. Open-source weights are due next week.
DeepSeek released V4-Flash-0731 on July 31, a retrained 284B MoE model that surpasses V4-Pro-Preview on all nine agentic benchmarks — including a 645% jump on DeepSWE — while keeping the price unchanged at $0.14/M input tokens.
Google has expanded Gemini Spark from the $99.99/month AI Ultra plan to all US Google AI Pro subscribers at $19.99/month. Spark is an autonomous agent that runs 24/7 on Google Cloud VMs, handling tasks across Gmail, Docs, Calendar, and 30+ connected apps.
Anthropic recommends giving each coding agent its own git worktree, but an analysis argues this pattern conflicts with the runtime infrastructure many teams already run.
Meta entered the AI coding agent market on August 5 with Muse Code (beta), a terminal-based tool powered by the new Muse Spark 1.2 model. It runs persistent async background agents throughout a session and ships with a 1-million-token context window.
Meta has unveiled its own coding agent, entering the AI software-development tools market led by OpenAI and Anthropic, according to reports.
Alibaba has unveiled what it calls its most powerful AI model yet, signaling renewed momentum in China's frontier-model race.
Anthropic says it is starting in-country Claude inference in India alongside new local partnerships.
A research firm's analysis reportedly found DeepSeek's new AI model to be by far the cheapest to run among widely known models.
Google DeepMind unveiled Gemini Robotics 2, a three-model series enabling whole-body humanoid control, multi-step task reasoning, and on-device inference. The launch marks a major push to become the default platform layer for physical AI.
LG AI Research published K-EXAONE 2.0 on Hugging Face under the Apache 2.0 license — a 750B-parameter, 37B-active hybrid-attention MoE model supporting 10 languages and a 262K-token context window, the largest AI foundation model built in South Korea.
Microsoft's first native full-duplex voice AI model, MAI-Realtime, surfaced in an internal MAI Playground preview. It handles simultaneous speech and listening, supports 18 languages, and can call web search and other tools.
OpenAI revealed that its unreleased Astra model solved ten long-standing open problems in mathematics and theoretical computer science, each accompanied by machine-checked Lean 4 proofs published on GitHub.
OpenAI reportedly said its new AI model, 'Astra,' made breakthroughs on 10 math problems. The scope and independent verification remain unclear.
Chinese outlet 36Kr reports OpenAI has unveiled a next-gen model called 'Astra' alongside a 249-page paper, claiming progress on 10 major unsolved AI challenges. Details remain unverified independently.
OpenAI slashed GPT-5.6 Luna API pricing by 80% and Terra by 20% just three weeks after launch, as Chinese AI models reportedly captured 46% of US enterprise token usage on OpenRouter.
DeepSeek officially released V4-Flash-0731 on July 31, 2026. The retrained model outperforms the company's flagship Pro-Preview across all nine published agent and coding benchmarks at just $0.14 per million input tokens.
The 'Reset to Zero' AI newsletter flags the compounding cost of long-running autonomous agents, using a 'the meter's running' metaphor to question their economics.
OpenAI reportedly slipped a mention of 'Astra,' said to be its next AI model, into a blog post about math, according to Gizmodo.
SitePoint examines how Claude may embed low-visibility markers in requests, and what developers should watch for in their pipelines.
YouTuber Hank Green publicly apologized, saying his LLM usage is unhealthy — a rare high-profile flag on chatbot dopamine and overreliance.
Forbes argues OpenAI's steep API price cut could trigger a margin-eroding price war across the AI industry.
Google DeepMind unveiled Gemini Robotics 2, a model that can control humanoid robots — a step toward 'physical AGI,' though deploying AI in the real world carries risks.
Microsoft put its own AI models and developer tooling front and center, signaling its most open competition yet with partners OpenAI and Anthropic.
According to TechCrunch, Microsoft is leaning harder into its own AI efforts, making its competition with OpenAI and Anthropic more visible than ever.
Anthropic released Claude Opus 5 on July 24, offering intelligence close to the flagship Fable 5 at the same price as its predecessor, plus a new effort setting that lets users trade cost for capability.
Moonshot AI released the full 2.8-trillion-parameter Kimi K3 weights for free download on Hugging Face, making the world's largest open-weight model accessible to anyone with sufficient compute.
The Financial Times ran an experiment casting LLMs as Fed FOMC members deciding on interest rates — and the models reportedly leaned toward a hike. A look at using AI for macroeconomic judgment.
In a Computerworld Q&A, Nvidia's generative AI lead lays out why open-weight models are central to enterprise AI adoption.
Global contract research organization ICON announced a multi-year collaboration with Anthropic to apply frontier AI across clinical trials.
LiquidAI unveiled LFM2.5-Encoders, an encoder model family aimed at fast long-context inference on CPU — targeting embedding and retrieval workloads without a GPU.
Starting July 20, Anthropic makes Fable 5 a permanent feature of Max and Team Premium plans at 50% of weekly limits. Pro and Team Standard users lose included access and must buy usage credits at $10/$50 per million tokens.
China's Moonshot AI launched Kimi K3, a 2.8-trillion-parameter open-weight model that jumped from 18th to 1st place in the Frontend Code Arena, beating Claude Fable 5.
Mira Murati's Thinking Machines Lab launched Inkling, a 975B-parameter open-weight MoE model built in under nine months, designed as a customizable foundation rather than a finished product.
DeepSeek is retiring the deepseek-chat and deepseek-reasoner model identifiers on July 24 at 15:59 UTC. Any production code that calls these endpoints will start returning errors — the fix is a one-line model parameter update.
Chinese open-weight models have taken 41% of Hugging Face downloads and 61% of tokens processed on OpenRouter, with all six of the platform's most popular models coming from Chinese firms. Qwen has surpassed Meta's Llama as the most downloaded model in Hub history.
Google DeepMind is targeting July 17 for the release of Gemini 3.5 Pro following a complete architectural overhaul after engineers discovered structural failures in recursive tool-calling and SVG generation. All specs — including the reported 2M-token context window — remain unconfirmed by Google.
Google launched Africa's first Applied AI Lab on July 1 at the Accra AI Community Centre in Ghana, offering founders and researchers across the continent early access to Google DeepMind models, VC mentorship, and go-to-market support.
OpenAI released GPT-Live on July 8, a full-duplex voice model that can listen and speak at the same time. It replaces ChatGPT's default voice experience across all plans, with GPT-5.5 handling complex reasoning in the background.
Anthropic released Claude Reflect on July 9, a built-in analytics dashboard showing users their conversation patterns, topic history, and peak usage hours — with quiet-hours controls and break nudges.
Google's Gemini 3.5 Pro, delayed from its June release window, is now targeting July 17 for general availability after a full architectural rebuild — bringing a 2-million-token context window and Deep Think reasoning.
Google has made Gemini 3.5 Flash the default model for its Search AI Mode and Gemini app worldwide, shipping a generative UI that assembles interactive visuals, tables, and simulations in real time instead of returning a ranked list of links.
OpenAI announced that GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture using 64 parallel subagents in less than an hour—though independent peer review is still pending.
ByteDance launched Seedream 5.0 Pro on July 9 — a multimodal image model for professional creators that adds deep-thinking reasoning, interactive precision editing, native 2K output, and on-image text in over ten languages.
Meta launched Muse Spark 1.1, a multimodal agentic model with a 1M-token context window, via a paid API priced at $1.25 per million input tokens — roughly one-quarter of Anthropic and OpenAI rates.
OpenAI launched ChatGPT Work, a GPT-5.6 agent that autonomously turns a goal into finished documents, spreadsheets, and code, while merging Codex into a single free desktop app.
Meta's Superintelligence Chief Alexandr Wang told employees that Watermelon, Meta's next frontier model still in training, matches GPT-5.5 on internal benchmarks, using 10x the compute of its predecessor.
OpenAI said its new GPT-5.6 model family will remain the 'preferred model' for Microsoft Copilot 365, pushing back on speculation about a partnership breakup.
Anthropic launched a public beta of Claude for Government on July 7, bringing Claude Code and Cowork to federal agencies through a FedRAMP High environment, priced at $1 per agency through August.
OpenAI opened GPT-5.6 to all users on July 9, with three pricing tiers. Budget tier Luna ($1/$6 per million tokens) marks the first sub-GPT-4o-mini price for a frontier-class model.
xAI released Grok 4.5 on July 8, a 1.5T-parameter V9 model co-trained with Cursor for coding and agentic tasks, priced significantly below comparable frontier models.
Anthropic's post-ban grace period giving Pro/Max/Team subscribers up to 50% weekly usage with Fable 5 ends July 7. From July 8, access is usage-credits-only at $10/M input and $50/M output tokens.
Anthropic released Claude Sonnet 5 on June 30, offering reasoning and agentic capabilities close to Opus 4.8 at an introductory price of $2 per million input tokens.
OpenAI is launching GPT-5.6 Sol on Cerebras wafer-scale infrastructure in July at up to 750 tokens per second — roughly 10× faster than standard GPU inference — with access initially limited to select partners.
ICML 2026, opening in Seoul with a record 23,918 submissions, desk-rejected 497 papers after detecting 398 reviewers who used LLMs to write their reviews via a hidden-watermark technique.
Anthropic announced an internal preclinical drug discovery program targeting neglected diseases that large pharma skips for economic reasons, unveiled alongside the Claude Science workbench on June 30.
Anthropic launched Claude Sonnet 5 on June 30, its most agentic mid-tier model yet, delivering performance close to Opus 4.8 for autonomous agent tasks at significantly lower cost.
xAI has quietly deployed Grok 4.5, its 1.5-trillion-parameter V9 model trained on Cursor coding data, to employees at SpaceX and Tesla. Elon Musk claims it matches or beats Claude Opus, though no independent benchmarks exist.
Chinese tech giant Meituan released LongCat-2.0 under an MIT license on June 30: a 1.6-trillion-parameter agentic coding model pre-trained exclusively on over 50,000 domestic Chinese ASICs — no Nvidia required. It is the first publicly confirmed frontier-scale model to complete its full training cycle on Chinese-made silicon.
Google has rolled out Gemini 3.5 Flash across its search engine, replacing traditional link lists with real-time AI-generated summaries — the company's biggest Search update in 25+ years.
Meta Superintelligence Labs chief Alexandr Wang told employees at an internal town hall that Watermelon—trained with ~10× more compute than its predecessor—has reached GPT-5.5-level performance on several key benchmarks.
Anthropic's Claude models are now generally available on Microsoft Azure AI Foundry as of June 29, running on NVIDIA GB300 Blackwell Ultra GPUs — cutting inference costs by ~50% versus A100 and generating tokens ~40% faster than H100 nodes.
Anthropic launched Claude Sonnet 5 on June 30, replacing Sonnet 4.6 as the default model for Free and Pro users. The model delivers near-Opus 4.8 performance at substantially lower cost, making capable agentic workflows significantly more affordable.
Google launched two new generative media models on June 30: Nano Banana 2 Lite for ultra-fast, low-cost images and Gemini Omni Flash for video generation and editing.
Anthropic released Claude Science in beta, a unified research environment with 60+ scientific databases and tools built in — framing the product as a workflow breakthrough rather than a new model.
Anthropic released Claude Sonnet 5 on June 30, making it the new default for Free and Pro plans with performance approaching Opus 4.8 and a significantly lower price point.
Anthropic's Frontier Red Team found that Claude Opus 4.7 completed robot dog programming challenges at least 10× faster than any human team, and on average 37× faster on tasks both parties attempted.
Google launched Gemini 2.5 Pro with Deep Think on June 22, posting 82.4% on GPQA Diamond and 89.8% on MMLU-Pro — beating Claude Fable 5 and GPT-5.5 to take the top spot for science and reasoning among publicly available models.
Google has confirmed Gemini 3.5 Pro will reach general availability in July 2026, featuring the largest deployed context window at 2 million tokens and a gated Deep Think reasoning mode.
OpenAI has previewed the GPT-5.6 family — Sol, Terra, and Luna — with the flagship Sol introducing an ultra mode that parallelizes subagents for complex tasks. Access is currently limited to ~20 government-vetted partners.
Google has capped Meta's access to its Gemini AI models due to insufficient computing capacity, delaying several of Meta's internal AI projects and forcing staff to use tokens more frugally.
OpenAI retired GPT-4.5 from ChatGPT on June 27, 2026, completing a 30-day sunset that migrates all conversations to GPT-5.5. The retirement marks the official end of OpenAI's dense-model generation before the GPT-5 family took over.
Anthropic released a Swift package integrating Claude into Apple's Foundation Models framework, enabling developers on iOS 27, macOS 27, and all other Apple platforms to route complex queries to Claude alongside the free on-device model.
Alibaba released HappyHorse 1.1 on June 23, pushing the model to second place on global video generation leaderboards. The timing is sharp: OpenAI's Sora has shut down and ByteDance's Seedance stalled internationally, leaving a wide gap for rivals to fill.
NVIDIA released Nemotron 3 Ultra, a 550B-parameter open-weights MoE model that ranks #1 among US open-weights models on Artificial Analysis's Intelligence Index, featuring a native 1M-token context window and full release of weights, training data, and recipes under a permissive commercial license.
Apple unveiled AFM 3 at WWDC 2026: five foundation models ranging from a 3B on-device model to a cloud Pro tier running on NVIDIA GPUs in Google Cloud, all while maintaining Private Cloud Compute privacy guarantees.
Jonas Adler and Alexander Pritzel — both AlphaFold veterans and Gemini contributors — are set to join Anthropic, Bloomberg reported June 24, marking the fourth and fifth senior departures from Google in a single week.
The European Commission selected the Domyn-led EUROPA consortium to develop a sovereign open-source AI model covering all 24 EU official languages at frontier scale — over 400 billion parameters.
Google's Gemini 3.5 Pro has slipped past its June 2026 general availability window, with release now expected in July. The flagship model was publicly promised for June by CEO Sundar Pichai at Google I/O.
OpenAI and Broadcom revealed Jalapeño, OpenAI's first custom ASIC built specifically for LLM inference. Developed in nine months with AI assistance, the chip targets large-scale deployment by late 2026 and is designed to reduce dependence on Nvidia.
Anthropic pulled back a planned billing change for Claude Agent SDK usage on June 15 — the same day it was scheduled to take effect — citing developer pushback and competitive dynamics ahead of its IPO.
Claude Fable 5 leaves included subscription tiers on June 23 and shifts to usage-based billing. Developers also face a new refusal response type and an always-on Adaptive Thinking mode that break assumptions built for older Claude models.
June 22 opens the primary Polymarket prediction window for GPT-5.6 at 83% probability. Leaked staging logs point to a 1.5-million-token context window and a redesigned reward audit pipeline targeting the alignment failures that plagued GPT-5.5.
xAI's Grok 4.3 is now generally available on Amazon Bedrock, adding a genuine third-party frontier alternative to GPT and Claude for AWS-native enterprise teams.
Prediction markets are pricing GPT-5.6 at 83% odds for a June 22–28 launch. OpenAI's chief scientist has reportedly confirmed the model is a meaningful step beyond GPT-5.5, with improvements in reasoning and agentic tasks.
Anthropic launched Claude Corps, a $150 million program placing 1,000 early-career AI fellows inside U.S. nonprofits. Fellows earn $85,000/year fully covered by Anthropic; the first 100-person cohort launches October 2026.
DeepSeek released V4 Pro in April preview — a 1.6 trillion total-parameter MoE model with a 1M-token context window, MIT license, and SWE-bench Verified score of 80.6%, priced at $0.87/M output tokens, roughly one-tenth of GPT-5.5.
Gemini 3.5 Pro, which Sundar Pichai promised at Google I/O would ship 'by next month,' remains in limited Vertex AI enterprise preview as of June 19 — with 11 days left in the window and prediction markets giving roughly 50–55% odds of a before-June-30 release.
Google has launched its first new smart speaker since 2020, placing Gemini AI at its core. The $99 device ships June 25 and targets multi-step conversational home control rather than simple voice commands.
Alibaba open-sourced Qwen3-Coder-Next, an ultra-sparse Mixture-of-Experts coding model with 80B total parameters but only 3B active at inference time, delivering up to 10× throughput versus dense equivalents and a 70.6% SWE-Bench score.
Andrej Karpathy, OpenAI co-founder and former head of Tesla AI, announced he is joining Anthropic to build and lead a new team using Claude itself to accelerate pre-training research.
NVIDIA unveiled Cosmos 3 at Computex 2026 — a fully open omni-model that handles text, image, video, sound, and action in a single system, compressing physical AI training cycles from months to days.
OpenAI released LifeSciBench, a 750-task benchmark built by 173 biotech and pharma scientists to measure whether AI can perform real life science research decisions — and even the strongest models clear only about a third of the tasks.
Google's June 2026 Pixel Drop delivers Android 17 alongside two generative AI features: Gemini Omni for video creation and Lyria 3 for AI-powered music composition.
Zhipu AI's Z.ai released GLM-5.2 on June 13 — a 744B-parameter MoE model with a functional 1-million-token context window, two thinking-effort levels, and MIT-licensed weights promised the following week.
MiniMax launched M3 on June 1 — an open-weight model using a new sparse attention architecture (MSA) to deliver a 1M-token context window, native multimodal input, and SWE-Bench Pro scores above GPT-5.5 at 5–10% of the cost.
Launched at Google I/O 2026, Gemini 3.5 Flash surpasses the previous-generation Pro on coding and agentic benchmarks while running 4x faster at 40% lower cost.
OpenAI replaced its default ChatGPT model with GPT-5.5 Instant, cutting hallucinated claims by 52.5% on high-stakes prompts. On June 9 the companion personalization feature — memory of past chats, files, and Gmail — expanded from paid tiers to all Free and Go users.
Apple unveiled a Gemini-powered Siri overhaul at WWDC 2026 alongside iOS 27 Extensions, a framework that lets users swap in Claude, ChatGPT, or Grok as their default AI across Apple Intelligence features.
Google unveiled Gemma 4, an Apache 2.0 open model built for advanced reasoning and agentic tasks, alongside TurboQuant — an ICLR 2026 algorithm that sharply cuts the memory overhead of KV caches in LLM inference.
Meta reversed its open-source strategy with Muse Spark, its first proprietary LLM built after a $14.3 billion, nine-month rebuild of its entire AI stack under Superintelligence Labs.
AIWire is launching a daily-updated section that distills the latest LLM news in both Korean and English.