Article

Grok 4.7: Choosing Between the API, Cursor and Grok Build

Compare Grok 4.7 access, API and Cursor pricing, Fast availability, caching and US routing before choosing a coding-agent pilot.

Editorial illustration for Grok 4.7: Choosing Between the API, Cursor and Grok Build: code brackets represent developer tools. Not documentary evidence.

xAI announced Grok 4.7 on September 21, 2026, giving teams building coding and knowledge-work agents a successor to Grok 4.6. The practical question is whether to evaluate it inside an existing application, in Cursor or through Grok Build. This guide compares their documented access and prices as of October 8.

The strongest case for a pilot is long-running work with checkable results, particularly for existing 4.6 users. Choose the access route first, then compare completed-task quality, time and cost. Published benchmarks justify investigation, not an automatic production switch.

What changes from Grok 4.6

xAI says 4.7 uses a larger base model and additional reinforcement learning focused on harder, multi-hour tasks. Its launch results report 71.0% on DeepSWE v1.1 versus 65.2% for 4.6, both at high effort. The same launch compares 46.3% for 4.7 on CursorBench 4.0 with 40.4% for 4.6, but that comparison uses xhigh versus high effort. These are developer-reported scores, not RohitAI measurements or an equal-cost guarantee.

The 500,000-token window and direct-API token rates are unchanged: xAI’s current pricing table lists the same bands for 4.6 and 4.7. The context-budgeting issue discussed in our Grok 4.6 analysis still applies. A move to 4.7 therefore needs a quality or workflow benefit, not merely a larger-window argument.

Choose the product you intend to evaluate

Grok 4.7 is available through the public API, Cursor and Grok Build, but those routes do not offer identical controls or billing. xAI’s developer guide distinguishes ordinary model access from the partner-only Fast option.

  • Direct API: use grok-4.7 in your own application or agent loop. It accepts text and images and produces text, with function calling and structured outputs. Choose this route when your own runtime and tools are what you need to evaluate.

  • Cursor: evaluate the model with Cursor’s agent tools. Cursor documents Fast as the default on Pro and higher plans. Its India-only Start plan fixes the model at medium effort and standard speed. Record both settings; plan defaults can otherwise make comparisons misleading.

  • Grok Build: choose this when you want xAI’s coding-agent runtime, which supports interactive terminal use and headless operation. A Build evaluation measures that runtime as well as the model; it is not interchangeable with a direct API call.

Grok 4.7 Fast serves the same model on faster infrastructure. It is available only through Cursor and Grok Build, is excluded from Build’s free tier and cannot be selected on the public xAI API. Faster generation alone does not establish faster completion after tool calls, retries and review.

Price the route, not just the model name

All rates below are USD per million tokens. xAI’s pricing table applies long-context rates to the whole request when prompt length reaches 200,000 tokens. Cursor’s model documentation instead starts its higher band only above 256,000 input tokens. Both support up to 500,000 tokens of context.

Route and speed

Input-token band

Input

Cached input

Output

Direct global API

Below 200,000

$2

$0.50

$6

Direct global API

200,000 or more

$4

$1

$12

Cursor standard

256,000 or fewer

$2

$0.50

$6

Cursor standard

Above 256,000

$4

$1

$12

Cursor Fast

256,000 or fewer

$4

$1

$12

Cursor Fast

Above 256,000

$6

$1.50

$18

The Cursor rows are on-demand token rates, not subscription bills: included usage and plan allowances affect what you actually pay. Do not transfer Cursor’s 256,000-token boundary to Grok Build; this review did not verify Build’s exact Fast cutoff.

Illustrative calculation, not a test: hold usage at 220,000 uncached input tokens and 10,000 total billed output tokens, including reasoning once. Exclude tools, priority processing, taxes, retries, images, compaction calls and subscription allowances.

  • Direct global API: (0.22 × $4) + (0.01 × $12) = $1.00.

  • Cursor standard: (0.22 × $2) + (0.01 × $6) = $0.50.

  • Cursor Fast: (0.22 × $4) + (0.01 × $12) = $1.00.

These are different rate-card outcomes for identical assumed counts, not evidence that Cursor halves a real task’s cost. Different runtimes can send different prompts and take different numbers of steps.

Caching changes the bill but not the input-length boundary: cached tokens still count toward xAI’s threshold. If 198,000 of the same 220,000 API input tokens are cache hits, the calculation becomes (0.022 × $4) + (0.198 × $1) + (0.01 × $12) = $0.406. The assumed 90% cache hit rate is not a prediction.

Check state handling and regional scope before a pilot

  • Preserve reasoning state. Grok 4.7’s Responses API always returns encrypted reasoning content. When managing history yourself, send the reasoning items back unchanged in the next input. This behavior is separate from the store setting and does not establish zero retention; Chat Completions is unchanged.

  • Configure cache affinity. Set prompt_cache_key on Responses, or x-grok-conv-id on Chat Completions, following xAI’s endpoint-specific guidance. Keep earlier messages stable and inspect the returned cached-token count; setting a key is not proof of a hit.

  • Check Batch eligibility. The Grok 4.7 model reference says Batch API is not supported. Do not budget a Batch workflow or discount from generic API documentation.

  • Bound long runs. The overview’s “no text output limit” wording is not an unlimited-context or unlimited-spend promise. Set application-level budgets, timeouts and stopping conditions.

For US processing, the regional endpoint documentation specifies a 10% token premium. That makes the $0.406 cached example $0.4466. Its guarantee covers API handling, inference, moderation and retained request data, but excludes Files, Collections, server-side tools and your network path to xAI. Inspect those services separately before treating an agent workflow as US-only.

Make the adoption decision on accepted results

For an existing 4.6 integration, start with held-out repository or document tasks and prespecified acceptance checks. Keep the runtime version, tools, effort level and budget fixed while comparing models. Then vary Fast or reasoning effort separately. Include short prompts and histories near the threshold of the route you will actually use.

Record failures, retries, reviewer minutes and elapsed time to an accepted result alongside token usage. For direct API calls, xAI exposes billed cost in usage.cost_in_usd_ticks: divide by 10,000,000,000 for dollars and sum across requests. That field includes token and server-side tool charges, unlike the token-only example above.

If the intended role is code review, measure useful findings, missed issues and false positives; a coding score does not settle review quality. The ReviewBench evaluation guide explains those distinctions.

Adopt 4.7 for the workloads where the pilot shows better accepted results within your time and cost limits. Keep the current model where it already meets those limits. If public-API Fast access or Batch support is a prerequisite, the documented offering does not meet that requirement today.

Methodology: AI-assisted reporting and analysis based on published xAI and Cursor sources, checked October 8, 2026. Cost examples are transparent rate-card arithmetic. No hands-on model or agent testing was performed; account-specific access, quotas and actual task performance remain unverified.