Grok 4.6 review + Imagine review

August 22, 2026

by Robert Viragh — questions/comments welcome at rviragh@gmail.com

tl;dr I assigned Grok a large task to complete autonomously and it built a product in 11 minutes.

Who this is for: anyone who is interested in using frontier AI’s.

There is a lot of competition in the AI model space. People can’t get enough of the frontier models, myself included. I usually use ChatGPT and Claude for most of my work. I actually pay a ton of money: I am on the Pro 20x on ChatGPT ($200/month), the 20x Max on Claude ($200/month), and also pay $30/month for Gemini. I’ve also tried Kimi K3 (a recent open weights model that is competitive with the frontier models) on API.

What I use models for: I code, run AI infrastructure (heterogenuous models running adversarial design review and code review while building and operating software), run agent harnesses, and I use them to research and answer my questions, look up information, and teach me things I’m curious about. I do this as a contractor under my company Taonexus. (I work remotely on U.S. time, if your company or team is hiring, I would be interested in a similar role as a full time employee, send me an email. I don’t require any visa sponsorship.)

What my screen looks like on a typical day: a bunch of terminals including Claude and ChatGPT, github repos, etc. I also communicate with my agents by email and Jira (project management). (click to expand)

Grok wasn’t previously part of my workflow. When I tried previous versions of Grok, the level of code that it produced wasn’t as good as Claude and ChatGPT.

I was intrigued by the Grok 4.6 announcement, which promised to change that: https://x.ai/news/grok-4-6

Grok 4.6 release page (click to expand)

I thought this was a very interesting capability, and I was eager to test it. It also showed very competitive benchmark scores:

Very competitive benchmark scores (click to expand)

I also thought this was an interesting capability:

Interesting long-range capability (click to expand)

Since the benchmark scores were so high, I wanted to try it out. I signed up to SuperGrok at the $30 subscription tier.

I previously demonstrated live use (for example when Google released the Gemma 4 offline model, I tested its 10 GB and 20 GB model on a Mac Mini, it got 550 views and a bunch of upvotes when I posted a comment about it). So I decided to also livestream it.

(A bit awkwardly, this was at 2:30 a.m. after my workday had finished, and my partner asked me to be quiet, so it’s actually a bit of an awkward livestream.)

Here is the livestream of using Grok 4.6 and Grok Build. (32 minutes). I start typing the prompt at 9min13s, press enter at 10min58s and it’s done and finished at 21min47s. This was one-shot, with zero further interaction from me. (Except for turning on auto-approve, as I hadn’t done that in the session yet.) That’s just under 11 minutes end-to-end.

One-shot prompt with substantial scope (click to expand)

Although the prompt is short, the scope of this actually quite substantial. I said it should build the comparison page for all the frontier models as of right now, and summarize all the offerings by everyone. Plus do research counts of the number of users for all of them. Plus make an email sign-up form. Here is the text form of the prompt (feel free to run it on a different model, I can put your results on this page as a comparison if you email me):

I’d like you to build a comparison page for all the frontier models as of right now (research the latest version) and summarize all the offerings by everyone and how they compare. Show market shares with a pie chart. Show how much investment is in each respective frontier model, or the market capitalization of its parent company, and how many users each one has as of today. Offer the user a signup to enter their email address to sign up for updates about the latest frontier models. Present this as an attractive page on localhost that I can review here. Don’t host it anywhere or post it anywhere.

What it build in 11 minutes

Here’s what it built.

Here are the screenshots:

Summary (click to expand)
The pie chart — hovering over the pie chart highlights the corresponding entry (click to expand)
valuations cards (click to expand)
user counts (click to expand)
detailed model cards comparisons, 1/2 (click to expand)
detailed model cards comparison, 2/2 (click to expand)
detailed scorecard comparison table (click to expand)
summary of everything (click to expand)

(it says “if you write software, Claude Opus 5” is its current recommendation, but for various reasons I use Opus 4.8 among the Opus models.)

user signup (on this screen I just tried the signup, it works.) (click to expand)

You can use the form on the page, it works. You can sign up if you like, if I get a few signups I’ll set up Grok to send out a weekly update of any developments or big releases among the frontier models. Grok performed really well on a news benchmark (beating every other model), so it’s great for news.

It’s also worth reading its methodology at the bottom:

It’s worth reading the fine print (click to expand)

It noticed that its research produced differing figures, tracking different things, and intelligently fused them into comparison metrics. This is also smart: “where sources conflict, this page prefers company disclosures, then audited filings, then named research firms.” That is a very intelligent way to organize sources, and helps explain why the model does so well on a news benchmark.

Conclusion

Given the large scope of research involved, it completed the work at a blazing fast speed. I give it a 9.5/10 on design quality as the kerning on the text it produces is very tight, it needs a bit of space to breathe.

these letters are really tight
1B is connected like a ligature

If I had to iterate on it, I would tell it to add some breathing room to the spacing of the font in the titles and headings.

Conclusion

It truly lives up to the promise that made me sign up: “turning a broad product idea into a working first version” by “researching unfamiliar domains, structuring the application, implementing the core interactions”, and “Grok 4.6 produces stronger first passes on visual and interactive projects…given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop.”

Bonus: access to Imagine video generation

The subscription I signed up for (SuperGrok at $30/month) allows generation of 15 second clips of 720p video in Grok Imagine. I tried it out by giving it a complete script based on a famous cartoon. I thought it did very well, I posted the results on X here: https://x.com/robertviraghhu/status/2090857565286522965

It followed the script exactly. If you need 15 seconds of video, SuperGrok can do it for you. Generation took 1 minute and 20 seconds.

If you’re hiring

What is your team working on, and what are your biggest challenges?
AI is becoming more and more capable, and I would love to help work with you operating, designing, orchestrating, and overseeing agent fleets to help you achieve your team’s objectives and continue iterating and shipping. I currently work remotely on U.S. time but I’m also interested on-site or hybrid in the Bay Area or wherever your team is located. If you’re working on something cool in AI, agent orchestration, research or development, let me know what you’re working on and let’s have a conversation on how I could help your team achieve your objectives. reach out to me at: rviragh@gmail.com or on linkedin.