allnewscastallnewscast
Breaking News
AI & Tech

AInews: xAI Unveils Grok 4.7 With Expanded Reasoning for Developers

Nic Reeve6 min read
AInews: xAI Unveils Grok 4.7 With Expanded Reasoning for Developers

AInews: xAI introduced Grok 4.7 on September 21, 2026, positioning the model as its strongest system yet for software development, agentic tasks and professional knowledge work. The release is available through Cursor, Grok Build, the Grok API, third-party coding tools, model routers and cloud platforms, according to xAI’s launch announcement.

What did xAI release on September 21?

xAI released Grok 4.7 as a new flagship model for work that requires extended reasoning, code generation and research. The company says the system is built on a larger base model than Grok 4.6 and received further reinforcement learning on difficult, long-running tasks.

  • Launch date: September 21, 2026, according to xAI.
  • Primary uses: coding, agentic workflows and knowledge work, according to xAI.
  • Model identifier: grok-4.7 on the xAI API, according to xAI’s release information.
  • Inputs and outputs: text and images in, with text output, according to xAI’s published model details.

The launch followed several days of speculation about a new Grok version. Reports published before the announcement pointed to possible appearances in cloud-service quotas, but those reports did not establish a public release. xAI’s September 21 announcement supplied the first direct confirmation.

What capabilities does Grok 4.7 claim?

Grok 4.7 is designed to stay engaged with complex tasks for longer, check its own work and handle large amounts of context. Those claims come from xAI and describe the company’s intended use case rather than an independent performance verdict.

  • Context window: 500,000 tokens, according to xAI’s model information.
  • Reasoning controls: low, medium, high and xhigh settings, according to xAI’s release documentation.
  • Default reasoning level: high, according to release details summarized by independent AI industry publications.
  • Tools: function calling, web search, X search and code execution, according to xAI’s API documentation.

The model’s focus is practical rather than limited to chat. A coding agent can use tool calls, inspect files, revise code and run tests. A research workflow can retain more material in one session. The 500,000-token window also gives developers room to pass large repositories, technical documents or multi-step task histories without splitting every request.

How did the new model perform in published tests?

Early benchmark figures come mainly from xAI’s own launch materials and should be read as company-reported results. Independent coverage has repeated the figures, but the available reports do not establish that every test used identical prompts, tools or evaluation rules across competing systems.

  • CursorBench 4.0: 46.3%, according to xAI’s reported benchmark results.
  • DeepSWE v1.1: 71.0%, according to xAI’s reported benchmark results.
  • EEBench: 64.0%, according to xAI’s reported benchmark results.
  • HealthBench Professional: 56.7%, according to xAI’s reported benchmark results.
  • Harvey Legal Agent Benchmark: 19.6%, according to xAI’s reported benchmark results.

These tests cover different abilities. Coding benchmarks measure software-engineering performance, while health and legal evaluations examine professional question answering or agent behavior. A single percentage cannot describe the model’s performance across every use case.

Users will also need to distinguish benchmark scores from production reliability. A system can produce a strong result under a controlled test and still require human review when it changes a live codebase, handles sensitive records or makes decisions in regulated fields.

Where can customers use Grok 4.7?

Grok 4.7 launched across developer products rather than as a single consumer-only update. xAI says customers can access it through Cursor and Grok Build, while API users can connect it to applications and coding agents.

  • Cursor: available at launch, according to xAI and independent release coverage.
  • Grok Build: available at launch, according to xAI.
  • xAI API: available under the grok-4.7 name, according to xAI.
  • Other routes: third-party coding harnesses, model routers and cloud platforms, according to xAI.
  • GitHub Copilot: independent release reports said the model was added for Pro, Pro+, Max, Business and Enterprise plans on launch day.

Availability can vary by product, plan and region. API access also depends on account approval, endpoint support and a developer’s chosen integration. A model appearing in a routing service does not necessarily mean that every customer receives the same limits, latency or tool access.

How much does Grok 4.7 cost?

xAI priced the standard API model at $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens, according to the company’s published pricing. Longer prompts use a higher price tier.

  • Input: $2 per million tokens below the 200,000-token prompt threshold, according to xAI.
  • Cached input: $0.50 per million tokens below that threshold, according to release documentation.
  • Output: $6 per million tokens below the threshold, according to xAI.
  • Longer prompts: $4 input, $1 cached input and $12 output per million tokens above 200,000 tokens, according to xAI’s listed rates.

Token pricing is only one part of the bill. Applications that run repeated tool calls, web searches or code execution can consume more tokens and generate extra service costs. Developers also need to account for storage, monitoring and the cost of human review.

What changed from Grok 4.6?

Grok 4.7 keeps the 500,000-token context window and the standard $2 input and $6 output rates associated with Grok 4.6 for shorter prompts, according to release reports. xAI instead presents the upgrade as a capability improvement aimed at harder tasks, stronger verification and longer-running work.

  • Model scale: xAI describes Grok 4.7 as using a larger base model than Grok 4.6.
  • Training: xAI says the newer system received extended reinforcement learning on harder and longer tasks.
  • Self-checking: xAI says Grok 4.7 verifies its outputs more reliably.
  • Pricing: the standard short-prompt API rates remain $2 per million input tokens and $6 per million output tokens, according to xAI.

The company also lists a faster version called Grok 4.7 Fast. Independent release coverage reported that the faster variant is available through Cursor and Grok Build at twice the standard token rates, while it was not listed as a public xAI API model at launch.

What should developers watch next?

The next test will come from real deployments. Developers will assess whether the model’s longer context and self-checking reduce debugging time, whether higher reasoning settings justify their latency and whether the system behaves consistently across different coding tools and cloud services.

  • Independent benchmark replication could clarify how Grok 4.7 compares with rival systems.
  • API documentation and regional rollout details will determine who can access each feature.
  • Enterprise users will examine privacy, logging, retention and permission controls before wider deployment.
  • Developers will compare the standard and Fast versions on cost, speed and task accuracy.

Grok 4.7 arrives as AI companies compete for developer workloads that involve more than text generation. The product’s success will depend on measurable software outcomes, predictable operating costs and the ability to keep human oversight in the loop when an automated agent changes code or handles sensitive professional work.

Sources

  1. 1.x.ai
  2. 2.releasebot.io
  3. 3.sqmagazine.co.uk
  4. 4.aiweekly.co
  5. 5.cellcog.ai
  6. 6.scriptbyai.com
  7. 7.ai-tldr.dev
  8. 8.kie.ai
  9. 9.orcarouter.ai
  10. 10.mungomash.com
  11. 11.layer3labs.io
  12. 12.iweaver.ai
  13. 13.x.com
  14. 14.cellcog.ai
  15. 15.ai-tldr.dev

Read more →

Related Articles

AInews: xAI Unveils Grok 4.7 With Expanded Reasoning for Developers
AI & Tech

AInews: xAI Unveils Grok 4.7 With Expanded Reasoning for Developers

AInews: xAI introduced Grok 4.7 on September 21, 2026, positioning the model as its strongest system yet for software development, agentic tasks and professional knowledge work. The release is available through Cursor, Grok Build, the Grok API, third-party coding tools, model routers and cloud platforms, according to xAI’s launch announcement. What did xAI release on September 21? xAI released Grok 4.7 as a new flagship model for work that requires extended reasoning, code generation and research. The company says the system is built on a larger base model than Grok 4.6 and received further reinforcement learning on difficult, long-running tasks. Launch date: September 21, 2026, according to xAI. Primary uses: coding, agentic workflows and knowledge work, according to xAI. Model identifier: grok-4.7 on the xAI API, according to xAI’s release information. Inputs and outputs: text and images in, with text output, according to xAI’s published model details. The launch followed several days of speculation about a new Grok version. Reports published before the announcement pointed to possible appearances in cloud-service quotas, but those reports did not establish a public release. xAI’s September 21 announcement supplied the first direct confirmation. What capabilities does Grok 4.7 claim? Grok 4.7 is designed to stay engaged with complex tasks for longer, check its own work and handle large amounts of context. Those claims come from xAI and describe the company’s intended use case rather than an independent performance verdict. Context window: 500,000 tokens, according to xAI’s model information. Reasoning controls: low, medium, high and xhigh settings, according to xAI’s release documentation. Default reasoning level: high, according to release details summarized by independent AI industry publications. Tools: function calling, web search, X search and code execution, according to xAI’s API documentation. The model’s focus is practical rather than limited to chat. A coding agent can use tool calls, inspect files, revise code and run tests. A research workflow can retain more material in one session. The 500,000-token window also gives developers room to pass large repositories, technical documents or multi-step task histories without splitting every request. How did the new model perform in published tests? Early benchmark figures come mainly from xAI’s own launch materials and should be read as company-reported results. Independent coverage has repeated the figures, but the available reports do not establish that every test used identical prompts, tools or evaluation rules across competing systems. CursorBench 4.0: 46.3%, according to xAI’s reported benchmark results. DeepSWE v1.1: 71.0%, according to xAI’s reported benchmark results. EEBench: 64.0%, according to xAI’s reported benchmark results. HealthBench Professional: 56.7%, according to xAI’s reported benchmark results. Harvey Legal Agent Benchmark: 19.6%, according to xAI’s reported benchmark results. These tests cover different abilities. Coding benchmarks measure software-engineering performance, while health and legal evaluations examine professional question answering or agent behavior. A single percentage cannot describe the model’s performance across every use case. Users will also need to distinguish benchmark scores from production reliability. A system can produce a strong result under a controlled test and still require human review when it changes a live codebase, handles sensitive records or makes decisions in regulated fields. Where can customers use Grok 4.7? Grok 4.7 launched across developer products rather than as a single consumer-only update. xAI says customers can access it through Cursor and Grok Build, while API users can connect it to applications and coding agents. Cursor: available at launch, according to xAI and independent release coverage. Grok Build: available at launch, according to xAI. xAI API: available under the grok-4.7 name, according to xAI. Other routes: third-party coding harnesses, model routers and cloud platforms, according to xAI. GitHub Copilot: independent release reports said the model was added for Pro, Pro+, Max, Business and Enterprise plans on launch day. Availability can vary by product, plan and region. API access also depends on account approval, endpoint support and a developer’s chosen integration. A model appearing in a routing service does not necessarily mean that every customer receives the same limits, latency or tool access. How much does Grok 4.7 cost? xAI priced the standard API model at $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens, according to the company’s published pricing. Longer prompts use a higher price tier. Input: $2 per million tokens below the 200,000-token prompt threshold, according to xAI. Cached input: $0.50 per million tokens below that threshold, according to release documentation. Output: $6 per million tokens below the threshold, according to xAI. Longer prompts: $4 input, $1 cached input and $12 output per million tokens above 200,000 tokens, according to xAI’s listed rates. Token pricing is only one part of the bill. Applications that run repeated tool calls, web searches or code execution can consume more tokens and generate extra service costs. Developers also need to account for storage, monitoring and the cost of human review. What changed from Grok 4.6? Grok 4.7 keeps the 500,000-token context window and the standard $2 input and $6 output rates associated with Grok 4.6 for shorter prompts, according to release reports. xAI instead presents the upgrade as a capability improvement aimed at harder tasks, stronger verification and longer-running work. Model scale: xAI describes Grok 4.7 as using a larger base model than Grok 4.6. Training: xAI says the newer system received extended reinforcement learning on harder and longer tasks. Self-checking: xAI says Grok 4.7 verifies its outputs more reliably. Pricing: the standard short-prompt API rates remain $2 per million input tokens and $6 per million output tokens, according to xAI. The company also lists a faster version called Grok 4.7 Fast. Independent release coverage reported that the faster variant is available through Cursor and Grok Build at twice the standard token rates, while it was not listed as a public xAI API model at launch. What should developers watch next? The next test will come from real deployments. Developers will assess whether the model’s longer context and self-checking reduce debugging time, whether higher reasoning settings justify their latency and whether the system behaves consistently across different coding tools and cloud services. Independent benchmark replication could clarify how Grok 4.7 compares with rival systems. API documentation and regional rollout details will determine who can access each feature. Enterprise users will examine privacy, logging, retention and permission controls before wider deployment. Developers will compare the standard and Fast versions on cost, speed and task accuracy. Grok 4.7 arrives as AI companies compete for developer workloads that involve more than text generation. The product’s success will depend on measurable software outcomes, predictable operating costs and the ability to keep human oversight in the loop when an automated agent changes code or handles sensitive professional work.

Nic Reeve·
Claude 4.8 Leak and Gemini 3.5 in Arena Shake Up the AI Model Race
AI & Tech

Claude 4.8 Leak and Gemini 3.5 in Arena Shake Up the AI Model Race

A major leak involving Anthropic’s unreleased Claude Sonnet 4.8 , fresh speculation around a new Claude “Cardinal” model family, and the quiet arrival of Google’s Gemini 3.5 variants in the popular LMSYS Arena benchmark have turned this week into a flashpoint for AI watchers, analysts, and creators following channels like Jaylin Williams’ AI news series. Claude Sonnet 4.8: What the Leak Really Reveals The story of Claude Sonnet 4.8 begins with a packaging mistake in Anthropic’s @anthropic-ai/claude-code npm library. Developers discovered that a 59.8 MB source‑map file had been accidentally published as part of a March 31, 2026 update, exposing roughly 512,000 lines of internal TypeScript and 1,900+ source files tied to the Claude Code product. Although no customer data, credentials, or live systems were compromised, the debug bundle included internal references that were never meant to be public. Among those references was a string for “sonnet-4-8” , listed in an internal “forbidden strings” or Undercover Mode filter intended to block engineers from accidentally mentioning unreleased model versions in logs, UI text, or commit messages. The same list reportedly included “opus-4-7” and codenames like “mythos” , hinting at a broader roadmap for Anthropic’s flagship Claude family. Crucially, what leaked was infrastructure code and configuration , not a model checkpoint or weights. There was no public model card, no API documentation for a Sonnet 4.8 endpoint, and no benchmark tables. That means the only hard fact confirmed by the leak is that Anthropic uses a Sonnet 4.8 version string internally in its tooling, and that the company is at least planning or testing a new generation of the mid‑tier Sonnet line. Nonetheless, the episode sparked intense speculation. Some posts circulating in the AI community claimed improvements such as a double‑digit boost on coding benchmarks, large jumps in vision accuracy, and new background “agent” capabilities for longer‑running tasks. While these claims appear to be based on references in the debug code and extrapolation from recent Claude 4.x releases, none of it has been confirmed by Anthropic. As of mid‑August 2026, there is still no official release of Claude Sonnet 4.8 via the Anthropic API, Amazon Bedrock, or Google Cloud’s Vertex AI. Anthropic has characterized the event as a human packaging error , asked for the removal of thousands of mirrored copies of the bundle from public repositories, and has not committed publicly to shipping a model under the Sonnet 4.8 label. Anthropic’s Model Codenames: Cardinal, Capybara, and Beyond The same discussion around Sonnet 4.8 has drawn attention to Anthropic’s growing web of internal codenames for its Claude models. Earlier analyses of the leaked Claude Code source have identified names such as Fennec (associated with an Opus 4.6‑class model), Capybara (linked to an experimental tier reportedly positioned above Opus in capability), and Numbat for models still in testing. In this context, community chatter about a line tentatively labeled Claude “Cardinal” has intensified. While details remain sparse, commentators describe Cardinal as a potential new family or sub‑tier that could sit between existing Sonnet and Opus offerings, or as an internal branch focused on tools, coding, and persistent agents. At this stage, Cardinal appears more as an inferred codename and roadmap hint than a shipping product with a public model card. Anthropic’s deliberate silence reinforces a pattern the company has followed in previous cycles: internal version strings and codenames often appear in tooling and leaks months before any formal announcement. The presence of names like Sonnet 4.8 or Cardinal in code does not guarantee that these models will launch under those exact labels, or even that all of them will reach public release. Gemini 3.5 Steps Into the Arena While Anthropic grapples with the fallout from its source‑map leak, Google’s latest models are making waves in a very different way: by showing up in LMSYS’s Chatbot Arena , the crowdsourced benchmark that pits large language models against each other in blind, head‑to‑head comparisons. Over recent weeks, new variants labeled along the lines of Gemini 3.5 have appeared on the Arena leaderboard. Though Arena typically uses anonymized identifiers for models in active blind tests, enough metadata and performance trends have emerged for observers to tie several strong‑performing entrants to Google’s newest Gemini generation. Early community impressions suggest that Gemini 3.5 maintains or improves on Gemini 1.5’s long‑context and multimodal strengths, while focusing on tighter instruction‑following and better coding performance. In many blind Arena matchups, users report that the 3.5‑class models feel more responsive for everyday chat and reasoning tasks, with competitive results against top‑end systems from Anthropic and OpenAI. Because Chatbot Arena relies on voluntary, crowdsourced votes, its rankings do not carry the same weight as formal academic benchmarks. However, the leaderboard has become an important real‑world signal of how models behave in the wild, capturing qualitative factors such as style, clarity, and robustness that are harder to summarize in a single numeric score. How Creators Are Covering the Shifts The rapid sequence of developments—leaks, codenames, and new benchmark entries—has given AI‑focused creators ample material. Among them is Jaylin Williams , whose AI news content (including the episode referenced in the Mshale listing) aggregates stories such as the Claude Sonnet 4.8 leak , the rumored Claude Cardinal line, and the arrival of Gemini 3.5 in Arena into digestible updates for developers and enthusiasts. In these roundups, creators typically emphasize three themes: Escalating competition among frontier models, as Anthropic, Google, and OpenAI iterate at a rapid pace and use both official launches and quiet evaluations in public benchmarks to test capabilities. Opacity and leaks as recurring issues, with internal tools and debug artifacts becoming unexpected windows into company roadmaps long before formal communication. Practical impact on users , from developers wondering when they can actually access Sonnet 4.8‑class performance to businesses evaluating whether to build around Claude, Gemini, or a mix of providers. What to Watch Next Looking ahead, the key questions for users and observers are straightforward. Will Anthropic officially announce a Sonnet 4.8 or Cardinal model in the coming months, and if so, how will it be positioned against Opus and rival systems from Google and OpenAI? Will the capabilities hinted at in internal code—ranging from stronger coding and vision performance to more persistent agents—translate into accessible, production‑ready features? On Google’s side, all eyes are on how quickly the Gemini 3.5 line moves from Arena experiments and limited rollouts into broad availability across Google Cloud and consumer products. Any shift in pricing, context length, or fine‑tuning options could reshape how startups and enterprises choose between providers. For now, the landscape is marked by contrast: Anthropic’s unintended leak offers a glimpse into where Claude may be heading, while Google’s Gemini 3.5 seeks validation in open competition. Together, they signal an AI ecosystem where product roadmaps are increasingly visible—not just through press releases, but through code, codenames, and the collective judgment of users putting these systems to the test.

Nic Reeve·
Anthropic Gives Claude Cowork Shared Memory with Chat for Persistent Context
AI & Tech

Anthropic Gives Claude Cowork Shared Memory with Chat for Persistent Context

Anthropic is rolling out a major upgrade to its AI assistant, giving Claude Cowork the ability to seamlessly reuse information it learns in regular chat. The company has merged the memory systems behind Claude’s chat interface and its Cowork desktop agent, so details you share in one surface can now automatically be used in the other. One Shared Memory Across Chat and Cowork Previously, Claude’s long‑term memory was largely confined to chat sessions and was synthesized periodically, meaning it could take up to a day before information carried over into new conversations. Cowork, which runs complex, multistep jobs on a user’s desktop or in the cloud, relied on its own background memory file and prompt stitching to simulate continuity. With the August 25 update, Anthropic has combined these mechanisms into a single, shared memory system that serves both chat and Cowork. Anthropic describes the change simply: the same memory now powers both Claude chat and Cowork. When users hand a task to Cowork—such as drafting reports, updating spreadsheets, or coordinating project documents—the context Claude has accumulated over months of chats is immediately available. Likewise, any new facts or preferences learned during Cowork runs are written back into the shared memory and become available in subsequent chat sessions. Real‑Time Memory, Not Just End‑of‑Chat Summaries Another important shift is how Claude updates memory. Instead of waiting to summarize an entire conversation once it ends, Claude now adds topics to memory in real time as users chat. This means that if a user mentions that a project deadline moved to September, that update can be reflected in memory almost immediately and show up in the very next interaction—whether in chat or Cowork—without requiring a manual “remember this” command. Anthropic’s support materials explain that when Cowork runs in the cloud, what Claude remembers from previous chats is automatically available, and what emerges during Cowork tasks feeds back into chat memory. Behind the scenes, each Cowork prompt is assembled from the user’s immediate request, their global instructions, and a relevant slice of the shared memory, allowing the AI to behave as if it has persistent awareness of roles, projects, and preferences. What Users Gain: Less Repetition, More Continuity The practical effect for users is that they no longer need to repeatedly brief Claude on who they are, what they are working on, or how they like to work every time they switch between chat and Cowork. Anthropic and independent commentators highlight several common scenarios: Persistent project context: Ongoing details such as quarterly goals, client names, and current project status can be retained across weeks or months and recalled in both chat and Cowork. Stable roles and preferences: If a user identifies themselves as an investment analyst, a teacher, or a particular type of creator, Claude can remember that role and tailor responses accordingly, even when individual chats are short or focused on different tasks. Cross‑device consistency: The shared memory applies across web, desktop, and mobile experiences, so moving from a browser chat to the Cowork desktop agent no longer breaks context. Tech industry observers note that this update positions Claude more directly as an AI “teammate” that can track medium‑ and long‑term workstreams instead of acting purely as a session‑bound chatbot. Transparency and User Control Over Memory The shared memory system arrives alongside a push for greater user control. Anthropic now surfaces everything Claude remembers in a dedicated Topics view within memory settings, where users can inspect, edit, or delete individual entries. Memory is stored as discrete, categorized entries rather than a single opaque summary, making it easier to remove outdated or inaccurate information. Users can also pause memory or reset it entirely if they no longer wish Claude to retain prior context. In addition, Anthropic provides guidance on importing and exporting memory, so the information Claude stores about a user is not locked in and can in principle be backed up or moved. Handling Sensitive Topics Anthropic has emphasized that the system is designed to minimize the capture of highly sensitive information by default. Topics such as health data, beliefs, and other potentially sensitive categories are excluded from memory unless users explicitly opt in via an “Include sensitive topics in memory” setting. For business customers, team or enterprise administrators can centrally control whether memory is enabled at all, and may choose more restrictive policies depending on corporate governance requirements. External reporting indicates that memory generation is turned on by default for free, Pro, and Max plans, while Cowork itself is not available on free accounts. For organizations that want to keep different workstreams separated, Anthropic has indicated that the only way to maintain fully separate memories for chat and Cowork is to use different accounts, since the new system treats them as a single unified space. Availability and Limitations The new shared memory capability began rolling out on August 25, 2026, across Claude’s web, desktop, and mobile experiences, as well as Cowork running in the cloud. Earlier in the year, memory support was limited to chat surfaces, and some third‑party analyses noted that Cowork lacked access to that long‑term context. Anthropic’s latest release notes and help center now explicitly state that memory works across both chat and Cowork when the latter runs in the cloud environment. There are still technical constraints. Cowork’s use of memory depends on cloud execution rather than purely local processing, and incognito or memory‑disabled sessions remain stateless by design. As with other AI systems, Anthropic cautions that Claude’s memory is selective: it prioritizes high‑level preferences and recurring topics rather than storing every detail of every conversation. A Step Toward More Personalized AI Workflows By unifying memory between Claude chat and Cowork, Anthropic is betting that users will value a more personalized and continuous AI experience, particularly for complex, ongoing work. The update reduces friction for individuals juggling multiple projects and gives enterprises a clearer path to building AI‑augmented workflows that persist over time. At the same time, the company is attempting to balance convenience with privacy and security by giving users fine‑grained controls and limiting sensitive data retention by default.

Nic Reeve·