Grok 4.7 is SpaceXAI's most powerful model for coding and knowledge work. It is built to work longer on difficult tasks, to check its own work more carefully, and it comes with the company's best-calibrated safeguards to date. SpaceXAI states that it is served at the same price and speed as Grok 4.6 and that it is highly competitive in its class. The model is available today in Cursor and in Grok Build, and it can also be reached through the Grok API, third-party coding harnesses, and model routers and cloud platforms, so teams can adopt it inside the tools they already use.
The context for the release is the growing demand for models that can stay productive across long, multi-step jobs rather than only short prompts. SpaceXAI highlights CursorBench 4.0, a benchmark that stresses longer-running coding tasks, and notes that Grok 4.7 sits at the frontier in price-performance there. A second theme is professional knowledge work: in GDPval and AA Briefcase, the model is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves on Grok 4.6 on both benchmarks and performs comparably to other frontier models, addressing the gap between short-answer demos and the extended work real jobs require.
Grok 4.7 uses a new, larger base model compared with Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. As a result, the model is better at verifying its own work and at managing longer context, two capabilities that matter when a single task spans many steps. SpaceXAI also trained Grok 4.7 to natively understand the Grok Bot harness, which makes it better at conversational tasks and general knowledge work alongside its coding strengths. Together these changes describe a model tuned for endurance rather than one-shot answers.
Across benchmarks, SpaceXAI positions Grok 4.7 as strong in several distinct domains. On software engineering it scores 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1, with an asterisk marking a high-effort score, and 37.6% on Terminal-Bench 4.0. On electrical engineering it reaches 64.0% on EEBench. For professional knowledge work, it scores 1,657 on AA Briefcase v1.1 and an Elo score of 1,735 on GDPval. It also records 19.6% on the Harvey Legal Agent Benchmark and 56.7% on HealthBench Professional. The announcement compares these figures with Grok 4.6, GPT-5.6 Sol, and Fable 5.1, and notes that Grok 4.7 is better at creating documents and presentations.
Safety and cybersecurity receive a dedicated section in the announcement. Grok 4.7 was built with an entirely new safeguard stack, and SpaceXAI describes it as the strongest model it has tested on refusals and jailbreak resistance. In dual-use domains such as cybersecurity and biological work, the company says it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio's biosafety benchmark at 62.4%. On HackerBench v0.3, SpaceXAI's benchmark for risky and malicious cyber tasks, the model allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work, giving it the highest safety score on that benchmark. The company frames this balance as important for defenders, who need the model to be useful on legitimate security work.
SpaceXAI has also started giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research. This detail shows that the safeguard work is not only a matter of refusal behavior but also of controlled enablement, where trusted partners can use the model for defensive security research while risky dual-use prompts are largely rejected for everyone else. Combined with the stated jailbreak resistance and calibrated refusals, the release frames safety as a core part of the model's design rather than an afterthought layered on top.
The overall approach behind Grok 4.7 can be summarized from the announcement's own description: a larger base model, a longer reinforcement learning run on a harder task mix, and explicit training to verify its own work and handle longer context. Native understanding of the Grok Bot harness is another part of that methodology, aimed at conversational and general knowledge work. The commercial angle is equally deliberate, with the model served at the same price and speed as Grok 4.6 while the company reports frontier-level price-performance on longer-running coding tasks. For teams needing more throughput, SpaceXAI also serves a fast variant with twice the output speed at twice the price.
The stated benefits follow from those design choices. Because the model works longer on difficult tasks and checks its own work more carefully, users can delegate multi-step jobs with more confidence that intermediate steps will be verified. Because it manages longer context and understands the Grok Bot harness, it can carry conversational and knowledge-work tasks that span more material. SpaceXAI also reports improvements in creating documents and presentations, and in benchmarks modeled on the work of lawyers, nurses, and financial analysts. Finally, the price and speed parity with Grok 4.6 means the capability gains arrive without a corresponding increase in cost per token for standard usage.
Concrete scenarios appear throughout the announcement. Coding is first, both in Cursor and in Grok Build, where the model can be tried for free, and through third-party coding harnesses. Longer-running coding tasks are the specific focus of CursorBench 4.0, where the model is positioned at the frontier in price-performance. Multi-hour terminal work is represented by Terminal-Bench 4.0. Professional knowledge work includes the tasks measured by GDPval and AA Briefcase, described as work done by lawyers, nurses, and financial analysts, as well as the Harvey Legal Agent Benchmark and HealthBench Professional. Cybersecurity defense is another scenario, including invite-only red-team access for select partners. Electrical engineering is covered by EEBench.
Grok 4.7 targets developers and teams doing coding and knowledge work. It ships in Cursor and Grok Build on day one, and is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms. Pricing starts at $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the output speed for twice the price. A free trial is offered through Grok Build at x.ai/build, and the announcement also includes a CLI install command. SpaceXAI points users to its Console for creating an API key and to docs.x.ai for documentation.
In short, Grok 4.7 is presented as SpaceXAI's most capable model for coding and knowledge work, combining a larger base model and longer reinforcement learning with a new safeguard stack. It works longer on difficult tasks, verifies its own work more carefully, and is served at the same price and speed as Grok 4.6 while claiming a frontier position on price-performance for longer-running coding tasks. For developers and professionals who need a model that can stay on task, balance capability with safety, and fit into existing tools, that combination is the core value proposition.