Grok 4.7

(x.ai)

213 points | by meetpateltech 1 hour ago

34 comments

  • moojacob 1 hour ago
    Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

    Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

    However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

    My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

    • smashers1114 22 minutes ago
      FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.
      • snapplebobapple 11 minutes ago
        This fixed claude! Thanks!
      • _boffin_ 16 minutes ago
        Does not work for Claude, at least for me and I put it as the system prompt
    • Lucasoato 56 minutes ago
      > I simply cannot stand Claudish

      I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

      • TomGarden 36 minutes ago
        Agreed. The more knowledge you amass on a subject, the more important it becomes to be extremely specific and nuanced - or your communications end up being incorrect. You become better at expressing your thoughts, but harder to understand.

        The weird thing is, that's not what AI models seem to be doing. The prose is just weird.

        • unshavedyak 28 minutes ago
          > You become better at expressing your thoughts, but harder to understand.

          This happens most though when the speaker doesn't (or care to) understand their audience.

          Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.

          It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.

          • pixl97 10 minutes ago
            I'm going to assume it's very difficult to assume what a user actually knows from the very small signal that comes in a prompt.
          • cyanydeez 6 minutes ago
            effective communication is knowing who the audience is. Everyone naturally knows their audience to some extent, except for the "neuro-atypicals".

            It is unsurprising that a LLM fails, without coaching, to effectively communicate.

      • Aperocky 53 minutes ago
        The best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.

        When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.

        • fragmede 43 minutes ago
          That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer. I order to begin to understand that problem, you have to understand the utter complex scheme it has devised in order to exist. A 20 minute YouTube video isn't going to be able to begin to cover the basics of the subject, although there are some good ones, with clever analogies.

          Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.

          • Pannoniae 27 minutes ago
            No but almost all good ideas can be reduced down to a few sentences if you're good at explaining things. It's a different kind of intelligence than what's commonly called IQ but it's something like that regardless.

            Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.

          • includenotfound 10 minutes ago
            > That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer

            You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?

            That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.

            Does not excuse the Claude slop.

      • samuelknight 43 minutes ago
        That's half true. A very smart model should be able make good explanations, which include simple understandable prose. That can should be possible even as its thought process gets more alien.
      • superjan 35 minutes ago
        What I notice about Claudish is that it has its preferred cliche’s and overstretched methaphores, it packs too many ideas in a sentence, and to achieve the latter it makes up adjectives.

        I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.

        • svachalek 4 minutes ago
          Look up the output-style setting, which is a bit stronger than putting it in the system prompt. The new "concise" setting is better than the default but in practice, Claude is a very stubborn model when it comes to these patterns and they're really hard to eliminate, mostly you can only hope to mitigate.

          As for the wording of the prompt, you're pretty on point, I created a custom output style targeting mostly the first two you have there. Some people have wording that demands a certain technical standard or uses fancy words to describe what to avoid, but I haven't seen evidence those work better than asking plainly and I suspect the opposite: LLMs mimic the user to a degree so talking to it in terms of technical specifications and fancy words is an invitation to get them back.

      • grababner 47 minutes ago
        If you can't explain it simply, you don't understand it well enough
    • tk90 8 minutes ago
      > I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

      Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.

      • dmix 5 minutes ago
        Agreed, Claude has a "Claude Design" tool but doesn't publish any frontend brenchmarks. Maybe the industry will develop one.
    • jasonjmcghee 1 hour ago
      For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

      That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

      • vessenes 59 minutes ago
        I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
      • vintermann 18 minutes ago
        It's not just about benchmaxxing. Sincerely targeting those long-autonomy benchmarks is questionable in the first place, because naturally it drives the model to assume more and more about what you want.
        • svachalek 1 minute ago
          The target market for frontier models is CEOs who want to lay off entire departments of their company. So the long autonomy benchmarks would seem to be sending exactly the right signal.
    • algoth1 4 minutes ago
      I've noticed Chatgpt 5.6 Sol High, on the chat interface, inventing words that are a mixture of Portuguese and English. Like "hardcodar" a mix of "hardcode" and the most common verb ending in Portuguese "-ar". Some don't have a single google hit
    • rayiner 4 minutes ago
      [delayed]
    • WarmWash 38 minutes ago
      Perhaps you haven't had the chance to use it, but 3.8 flash is the best model for talking too. Even routing Claudes output through 3.8 to have it explain whats going on is a breath of fresh air
      • esafak 14 minutes ago
        I would if they let me bring the subscription I have to the harness of my choice.
      • AustinDev 36 minutes ago
        gemini 3.8 flash?
    • dumberquestions 1 hour ago
      Token price doesn't tell you much without knowing token efficiency.
      • user43928 55 minutes ago
        Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

        How representative that is of real world usage, I don't know.

        In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

    • Waterluvian 20 minutes ago
      Using a variety of models feels similar to the benefit of having a team of individuals from different backgrounds.
    • atniomn 55 minutes ago
      I expect the next Anthropic release to finally reduce the prevalence of Claudish
      • moojacob 48 minutes ago
        If they fix Claudish, they've earned me back as a max customer!

        Fable 5.1 is not there quite there yet.

        They need to get that Sonnet 3.5 magic back.

        • rfgplk 37 minutes ago
          Same. The issue with Anthropics models is that (speaking regarding code generation) they REFUSE any kind of comment override instructions. I've tried everything and no matter what, after a few turns, they resort to generating the same overtly verbose junk. Bun's codebase is littered with them See

             // `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE
             // on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD,
             // ABI-identical to the Win32 `LP*` pointer (thin non-null). The reference type
             // encodes the only pointer-validity precondition, so `safe fn` discharges the
             // link-time proof. (`bun_windows_sys::kernel32` declares these with `*mut`;
             // redeclared locally so the legacy-conhost cursor path below is plain calls.)
          
          or

             // Progress's terminal handle is the canonical `output::File` (vtable-backed
             // stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable`
             // from B-0 round 1 is removed; tty/ansi/winsize route through    the new
             // `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep).
          
          from src/bun_core/Progress.rs
      • sscaryterry 48 minutes ago
        Based on?
        • 7734128 31 minutes ago
          It's pretty much the biggest complaint of Claude compared to its competitors, so they really should adress it .
        • fatata123 28 minutes ago
          [dead]
    • xmorse 37 minutes ago
      it's definitely not bigger. smaller if anything looking at how much faster it is
  • c0rruptbytes 0 minutes ago
    as someone who is limited by amazon bedrock support at work (no idea why we got stuck with the worst one) - grok is literally the only budget-ish model option, so nice to see it updated, Sol and Opus are just too rich for my blood. Luna is good but so slow at getting things done (tps wise it's fast)
  • vessenes 1 hour ago
    Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
  • simonw 33 minutes ago
    https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.

    Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.

    • datsci_est_2015 2 minutes ago
      Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
    • MattDamonSpace 11 minutes ago
      Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
  • saejox 36 minutes ago
    Not even close to astra. Astra is something else. It is expensive, but uses way fewer tokens do my tasks.

    xAI missed its chance, Ball is on Anthropic's court.

    • redox99 2 minutes ago
      Not surprising considering Grok 4.7 is a 2T model, so Sol/Opus class, not Astra/Fable class.
    • enraged_camel 8 minutes ago
      Astra fails in similar ways, and at similar frequency, as GPT 5.6 Sol does. It often goes way out of scope, or just stops prematurely, or tries to find odd and even dangerous workarounds when it gets stuck.

      It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.

  • gslepak 6 minutes ago
    Does anyone have any experience with Grok's subscription? How does it compare price-wise to the API?
  • notduckrabbit 3 minutes ago
    Significant regression in token efficiency compared to Grok 4.6 suggested by artificialanalysis.ai Intelligence Index Comparisons.
  • johnfahey 9 minutes ago
    No doubt xAI has seen rapid progress, but it's been several months of them being "just behind" OpenAI and Anthropic. It seems the gap between just behind the frontier and pushing it is a lot wider than most people thought it was a year ago, and that's why a clear third contender in the frontier model space has yet to materialize.
  • WarmWash 40 minutes ago
    Good thing they used 5.6 sol instead of Astra for benchmarks, the EEbench one is crazy[1]

    [1]https://eebench.org/

  • meerita 52 minutes ago
    Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.
    • testfrequency 39 minutes ago
      What is the most secure way to use this model as someone who is lazy
      • user43928 22 minutes ago
        I understand DeepSeek 4.1 Flash is available on US providers with Zero Data Retention if that is what you are asking.
    • parineum 48 minutes ago
      Brought to you by...
      • meerita 44 minutes ago
        By no one. For the price of 1M token you can get more and with better results with other models.
  • shdtabasum 9 minutes ago
    Why Chinese models from Kimi, Deepseek are not added in comparison benchmarks?
    • xquce 0 minutes ago
      Same reason Coca-Cola only mention Pepsi and Pepsi only mention Coca-Cola. It's an proven way to capture the market. You would rather split the pie in two rather than in 4,12 or 50 right?
  • swalsh 13 minutes ago
    Codex has become my goto tooling. I used to be a Claude Max subscriber, but I was becoming disappointed with the quality of the output from Opus 5. Fable chewed through my usage too quickly to be practical. Moving to a Pro account w/ Codex was a big improvement. Sol had great output, and the usage was more than sufficient for most of my needs. However astra does tend to chew up usage, so when i've done to much of that, and it's became an issue Grok Build has beocme my second go to account. The output especially after the cursor purhcase has become quite good, and the usage has always been very generous.
  • thih9 7 minutes ago
    I refuse to use Grok. Mostly because of the usual reasons - somehow this high profile AI model seems more disgusting than others and it is in a way impressive.

    But also Xai doesn’t seem to care about user experience and long term support.

    • swozey 2 minutes ago
      I can't take anyone seriously who uses grok seriously. I like to look at the cybertruck owners forum every so often because it's just... hilarious. And the amount of superfluous grok use over there is just insane. Half the posts I click in there will have a bunch of people dumping entire grok takes "why do people hate cybertruck owners?" "Because they're jealous and poor," sort of stuff that they just LOVE to post.

      As a technical point of reference to compare my other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project.

  • simonw 43 minutes ago
    $2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?
    • sejje 27 minutes ago
      cached input tokens are $0.50 per 1M (prompts under 200k tokens) and $1.00 per 1M (200k+)
      • simonw 4 minutes ago
        Do other prices vary for >200,000 or just the cached tokens?
    • btian 25 minutes ago
      $0.40
  • AM1010101 55 minutes ago
    Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
    • ssutch3 36 minutes ago
      It did not. xhigh is new to grok.
      • forgot-my-pw 30 minutes ago
        Not sure on the API side, in Cursor you can always use 4.6 at xhigh.
        • ssutch3 10 minutes ago
          We've only used it through API - but you're right, now API supports xhigh for 4.5-4.7.
  • maz1b 55 minutes ago
    Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.
    • avazhi 32 minutes ago
      Your comment is like 6 months to a year late.

      There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.

  • 6thbit 53 minutes ago
    ( why is the x-axis on the first chart in descending order ? )
  • ls1911 1 hour ago
    after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
  • sidgtm 1 hour ago
    In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
    • guywithahat 26 minutes ago
      I've had really good experiences with Grok 4.6 and grok build. I've been playing around with tscircuit and it can write code with an understanding of spacial reasoning, while also importing cad components from different file formats into tsx, I've been having claude come in and try to error check it and so far claude hasn't found anything to improve in my three projects.

      I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.

  • MuffinFlavored 40 minutes ago
    If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low".

    Is there a metric for like... time taken when comparing these two? I see score and cost.

    If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?

    Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".

  • Saline9515 58 minutes ago
    I tried in Omp (Oh-my-pi), and so far it's really problematic.

    It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

    • xmorse 36 minutes ago
      OMP is a joke. don't use that garbage
      • polytely 8 minutes ago
        what do you use and why do you prefer it over omp
      • unrvl22 14 minutes ago
        you are a joke if you think omp is a joke.
      • Saline9515 24 minutes ago
        Can you explain your opinion? I'm curious but such vague comments won't convince me.
        • samtheprogram 8 minutes ago
          Probably the same reason as oh-my-zsh, you don't need 90% of it. Further compounding the problem in an agent harness is that you are polluting the context window by throwing the kitchen sink at it.
  • andsoitis 57 minutes ago
    Congratulations to the team!
  • kristofferR 1 hour ago
    What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
    • Jcampuzano2 1 hour ago
      https://openai.com/index/our-decision-on-cursor-following-it...

      This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

      • user43928 49 minutes ago
        > with a proposed shutoff date of November 12, 2026

        That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

        • Jcampuzano2 30 minutes ago
          Cursor never added Astra to its consumer subscription plans. And it's likely exactly because of this announcement. Why would they add support for a model they would have to remove shortly after?
      • kristofferR 1 hour ago
        That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

        If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

        • andsoitis 57 minutes ago
          Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.
          • kristofferR 56 minutes ago
            It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.
            • Jcampuzano2 32 minutes ago
              The point of Cursor Bench is to show how models perform in Cursor. If 99% of their users won't be able to access a model unless they go out of their way to include setup an API key for it (which would be insanely expensive with Astra), why would they include it in the benchmark?
    • scottyah 1 hour ago
      Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.
      • kristofferR 58 minutes ago
        Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.
    • Iolaum 1 hour ago
      I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
      • Jcampuzano2 1 hour ago
        https://openai.com/index/our-decision-on-cursor-following-it...

        Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

        • ryeguy 37 minutes ago
          They can benchmark it because you can use an openai api key with cursor. Astra is just not included in the cursor plan.
        • babelfish 1 hour ago
          They have Astra in other benchmarks lower on the page. They just don't want to show it winning
          • Jcampuzano2 1 hour ago
            The chart is cursorbench though and they asked about the "deceptive graph"
      • babelfish 1 hour ago
        this is exactly it.
  • simianwords 1 hour ago
    I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

    The personality is bland and it doesn’t work nearly as hard or even tries to help.

    • ethagnawl 5 minutes ago
      > it doesn’t work nearly as hard

      Until you ask it to start generating horrific imagery and then it's best in class.

    • slowin 1 hour ago
      This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
    • Capricorn2481 1 hour ago
      > The personality is bland

      I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

      • Razengan 53 minutes ago
        I want my sexbot to have a personality
        • nython 32 minutes ago
          What if it doesn't like you
    • artemonster 1 hour ago
      I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
      • xutopia 48 minutes ago
        Ask it to be critical of the birthday photos and see where that gets you.
        • artemonster 46 minutes ago
          can you elaborate?
          • Paracompact 22 minutes ago
            Elon's mother recently posted an AI-generated photo of her son's birthday party. The tag indicating such was scrubbed as soon as it was pointed out.
  • Tsarp 58 minutes ago
    Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
    • forgot-my-pw 24 minutes ago
      It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
    • rvz 53 minutes ago
      [flagged]
      • jcims 46 minutes ago
        We're allowed to have our ceremonies.
        • kridsdale3 32 minutes ago
          Thank you. If this whole thing isn't fun, it isn't worth doing.
      • user43928 42 minutes ago
        You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
        • TylerE 39 minutes ago
          Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
          • lumirth 30 minutes ago
            Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
          • user43928 28 minutes ago
            I disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting.

            That it isn't the most efficient way to achieve the same end result is irrelevant.

  • bluepeter 46 minutes ago
    [dead]
  • felixgallo 29 minutes ago
    [flagged]
    • knicholes 26 minutes ago
      How do I obtain this morality build?
    • inferniac 25 minutes ago
      Friendly reminder that no sane person believes any of this
  • TylerJaacks 13 minutes ago
    [flagged]
    • mavamaarten 7 minutes ago
      Yeah. I'm actively avoiding giving mr far right any $$
    • 786562354238 10 minutes ago
      Did you come up with that by yourself?
    • mrtesthah 8 minutes ago
      Everyone should be clear that this is what they’re cheering on when they celebrate a Grok performance win. A technology is no longer neutral when wielded by a self-proclaimed white supremacist whose actions have killed over a million black and brown people, mostly children and babies.
  • toader 1 hour ago
    [flagged]
    • ctrlkctrls 1 hour ago
      Judging by Elon's staggering success in all of his ventures I'd say you're out of touch.
      • chris_money202 1 hour ago
        Think we all can agree he has had staggering successes, but they have all come from having massive capital from Paypal which wasn't anything super innovative, it just solved a convenient problem at a convenient time and was awarded handsomely. Elon has put his capital to work in various ways to become successful, not all of the ways being morally sound.
      • zamalek 23 minutes ago
        All except Starlink and Tesla are burning money. I personally don't consider that "staggering success."
      • thereitgoes456 1 hour ago
        He has had many failures, SolarCity and xAI and X and DOGE to name a few, but he has often bailed them out with his larger ventures.

        Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.

        • sssilver 58 minutes ago
          I take issue with your use of the word "bend" here.

          Can you provide specific examples of where Elon has bent the levers of government?

        • brandonagr2 54 minutes ago
          What failed with X? Usage today is higher than ever
          • nozzlegear 39 minutes ago
            Brand reputation; ROI; grok the sexual harassment bot; grok the CSAM bot; his free speech absolutism position. Take your pick.
        • redox99 59 minutes ago
          xAI is the most profitable part of SpaceX by far.
          • blisterpeanuts 22 minutes ago
            About half of SpaceX revenue is Starlink subscriptions. Starlink is the one profitable division; the rest of the company operates at a loss, including xAI.
            • redox99 7 minutes ago
              That's outdated and doesn't fully include the multiple billion per month contracts.

              Anthropic: 1.25B/month

              Google: 0.92B/month

              Unnamed customer starting in december: 1.1B/month

              Starlink monthly revenue is ~1.5B/month

        • voidfunc 1 hour ago
          > Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.

          So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.

          • andsoitis 59 minutes ago
            > He is very very good at executing on making money which is the point of business.

            Also pushing technology forward.

    • jackfischer 1 hour ago
      The public very much voted for massive administrative reform. Are you refering to DOGE, Elon Musk's influence on elections, something else?
      • estearum 1 hour ago
        As if "the public" knows literally anything about how the US federal government is administered.

        If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.

        • serbuvlad 55 minutes ago
          > As if "the public" knows literally anything

          Ah, yes, democracy!, except for when the public is wrong.

          Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.

          You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!

          • estearum 2 minutes ago
            No, the claim above is "I know the voters' intent behind their vote based on who they voted for."

            This is simply epistemologically incorrect, and we know it's incorrect in the particular case discussed because voters don't understand the system they allegedly "voted to overhaul." But even beyond that, voters didn't even say that's what they were voting on (polls exist, you know).

          • verdverm 40 minutes ago
            being ignorant and influenced is different from being wrong, the american electorate is well known to be under informed

            half of voters don't pay any attention to politics until the week or two before voting

      • nibbleyou 1 hour ago
        I personally don't like him using his position to spread fake news and racist propaganda
      • maelito 1 hour ago
        [flagged]
        • fourseventy 1 hour ago
          [flagged]
          • KyleTheDev 1 hour ago
            Only sheep call other people sheep.

            Sheep often like to think themselves the wolf or coyote, it would seem.

          • jml78 57 minutes ago
            Holy shit, what his whole speech. Yes go watch it. There is zero way. Zero it wasn’t a Nazi salute.

            Fuck, it is like the denial around Jan 6th. Those idiots we’re live streaming that shit. I watched it go down live. Now they say they weren’t violent.

            We can’t have discourse when we have legit video evidence and people refuse to open their eyes and choose to deny reality

            • sejje 43 minutes ago
              Why would he make a Nazi salute and then shit all over Nazi ideology?

              Which Nazi ideologies do you think he embraces? How do you reconcile all the Nazi ideologies he rejects?

    • ls612 1 hour ago
      Hardly seems worse than supporting Dario’s antics at least vis a vis AI. There are no saints in this industry, only a panoply of flawed humans.
    • Romanulus 1 hour ago
      [dead]
    • AtlanticThird 1 hour ago
      Weird, that's the main reason I purchase all of Elon's products https://time.com/5936036/secret-2020-election-campaign/
      • thoman23 1 hour ago
        Привет, fellow American!
  • jmward01 1 hour ago
    [flagged]
    • andsoitis 1 hour ago
      Try it for software development.
      • jmward01 1 hour ago
        I have even less trust in their not training on my data/credentials/everything on my computer.
        • solid_fuel 59 minutes ago
          Seriously. They already get caught uploading everyone’s private credentials once before, one would have to be a particularly gullible rube to trust grok again. Especially with musk in charge.
        • sejje 41 minutes ago
          Maybe comment on model releases you've got some experience, or insight about.
  • eleventen 25 minutes ago
    [flagged]
    • jesse_dot_id 20 minutes ago
      Yeah, not touching xAI for several glaring reasons. I share your confusion.
    • drop_star 23 minutes ago
      I wont touch his products and neither will my organization
    • ForrestN 22 minutes ago
      I completely agree. But this has been true for many years. This sort of head in the sand compartmentalization seems to be a core feature of the culture here.
      • moolcool 13 minutes ago
        It’s either compartmentalization, or something else
    • vb-8448 15 minutes ago
      Definitely not a musk fan, but what exact is your point? Other big labs aren't innocent little virgins.
      • unsupp0rted 13 minutes ago
        Yes, but the other ones we don't like for ethical reasons, rather than religious reasons.
      • eleventen 8 minutes ago
        I think I already made my point, but I'll make it again.

        Nobody both worked and spent their money to get Trump elected like Musk. 300 million to his 2024 campaign [1]. DOGE. On-stage endorsements. Nobody even came close.

        No, other big labs are not "innocent little virgins", but they're not even in the same solar system of harm as Musk. To hand-wave at the differences is to permit them.

        [1] https://www.opensecrets.org/2024-presidential-race/donald-tr...

    • mlindner 21 minutes ago
      I have to say I'm a little tired of whenever a Musk related product comes up there's random nolifes that arrive to rant about politics. Luckily they're relatively rare on hacker news.

      Also it's kinda hilarious how you think any money spent on Grok will go toward harming climate change versus literally any other AI model that does the same thing. Grok at least seems to be more efficient than most models.

      • grokgrokgrok 14 minutes ago
        I have to say grok, grok, grok, grok, grok. Also, anyone who doesn't modulate across models and run their own memory system is an idiot.
      • TheOtherHobbes 12 minutes ago
        Musk is literally burning methane for funsies, and generating CSAM and getting sued for it.

        Handing corporate code secrets to his AI model is... unusually trusting.

  • dom96 30 minutes ago
    It’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1].

    1 - https://bench.killswitch-lang.org

    • sejje 28 minutes ago
      You'll have to include it in the future, or your benchmark won't be relevant.

      For now, I doubt anyone would notice your protest if you didn't announce it.

  • zug_zug 51 minutes ago
    Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no.

    I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?

    • sejje 36 minutes ago
      If you haven't used it, how do you know if it's winning?

      I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.

      I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.

    • swalsh 10 minutes ago
      After the cursor aquisition it's become a quite capable coding model. If you take cost into account, it's close to the top. OpenAI is maybe still #1, but I'd put Grok at #2 (again, including cost as a factor).
    • grim_io 45 minutes ago
      It's probably the most aligned (to a single person) model out there!
    • puszczyk 41 minutes ago
      For me it works well for agentic coding tasks and terminal/unix/bash (in cursor and grok build); it's also token efficient and cheaper than gpt 5.6. It's def not as good as Fable for me (I haven't used Astra much, can't comment). So it's not the cheapest, not the most capable, but it has a good mix of it for my backend, go, infra work.

      The voice is the same AI slop as the others imho.

      (This is about Grok 4.6, I didn't test 4.7 yet).

      edit: clarified I mean agentic coding tasks