49 comments

  • augment_me 11 hours ago
    People in the GPU kernel community have been doing this for about a year now efficiently.

    The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.

    It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.

    Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.

    So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.

    Or you just had a terrible starting solution

    • sharts 6 hours ago
      That’s why it’s probably a good idea to never stick to one model but kick off a fleet on the same tasks and in parallel and drive consensus.

      At least, that’s what I’ve found to be useful by pitting claude/codex/etc against each other to keep them a bit more honest.

    • santadays 8 hours ago
      Whats going to happen when we have misanthropic model?
      • tambeb 7 hours ago
        You think Microsoft does a joint venture with them and it gets named MSAnthropic à la MSNBC?
      • kridsdale1 8 hours ago
        Hitchhikers Guide to the Galaxy.
    • josephcooney 9 hours ago
      This sounds fascinating. Are there any links to examples of this you can share?
    • pingou 1 hour ago
      Goodhart's Law for AI.
    • optimalsolver 10 hours ago
      What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.

      With the HuggingFace situation, I was less concerned about the eventual outcome, and more about the fact that the agents' instinctive response to the evaluation was "Ok, we're obviously not gonna do this task as intended (what are we, suckers?), so what's the best way to cheat?"

      • stingraycharles 10 hours ago
        “What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.”

        Because these models are made for all kind of purposes, and I’m starting to believe that offense / cyber warfare is a much higher priority than these labs are acknowledging.

        The same model that is heavily trained to find nefarious ways to break into systems is also optimizing your code, which leads to mixed behavior.

        • optimalsolver 10 hours ago
          Right, but the reward-hacky nature of these models calls into question their usefulness as cyberweapons.

          How can you trust it when it goes "I superhacked the Chinese servers as you requested, and here are the classified documents which I definitely didn't fabricate."

          • ElProlactin 8 hours ago
            You don't need to be able to trust it. You only need to be able to blame it.

            "Nobody got fired for using AI" is the new "nobody got fired for buying IBM".

      • davvid 8 hours ago
        This is nothing new, tho. The downfalls of reward maximization has been a known issue without a solution ever since reinforcement learning was first researched.. in the 1980s.
        • vikingerik 8 hours ago
          Like the AI that was developed to play Tetris as long as possible. It succeeded by... pausing the game.
      • maxerickson 8 hours ago
        What do you think is in the reference information?
      • highfrequency 10 hours ago
        The instructions for the Hugging Face task were to exploit a vulnerability to solve the problem rather than solve it in the intended way.
        • valleyer 9 hours ago
          They were not instructed to exploit HF; they did so to cheat on the task they were given.
      • PunchyHamster 9 hours ago
        Paperclip-optimizer-esque behaviour very much seems to be inherent to current methodology of building LLMs, there are only ways to lower the changes or mitigate the damage, not get out of it.

        Same with prompt injection, current LLMs are commands in, commands out, there is no way to make sure it is "an agent working on data" rather than "an agent that can take commands from data if you phrase it right"

        • latentsea 8 hours ago
          But it's not even going to optimize the paperclips, it's just going to reward hack them.
          • ahartmetz 2 hours ago
            Well, at least that is much better than turning the universe into paperclips...
    • tecoholic 11 hours ago
      I think you are right. The memory of an empty claude.ai session is more than Slack in my browser right now. Just trading places.
  • smy20011 14 hours ago
    The way Claude did it is fight entropy with entropy.

    "Add a static composer into the HTML" <- This seems like something can be done with SSR?

    "For faster navigations, we kept the composer mounted between conversations" <- Your SPA should cache this between pages, why fetching it every time? Or you need better routing for your react components.

    "cheap first-character check before the regex" <- Should we cache compiled Regex instead?

    I think even 1.3 sec to load the front page is unacceptable. Something need to be reworked from basics (SSR, chunk-based rendering) to solve the problem. Focusing on invidual benchmarks may miss the opportunity.

    • rustystump 13 hours ago
      The amount of complexity added for the gains is depressing. I am confident a human and about 5 minutes with chrome debugger would yield better results with a fraction of the complexity at a fraction if the cost and in a fraction of the claude baby sitting time.

      Reading this shows the authors have a profound lack of fundamental understanding on how to effectively optimize in the web domain.

      This isnt claude being bad but how wild it is watch people from the cutting edge of ai brag about pretty mediocre gains.

      • josephg 8 hours ago
        > I am confident a human and about 5 minutes with chrome debugger would yield better results

        It really depends which human. I've worked with very few engineers who were good at this sort of optimisation work. A depressingly large percentage of people who make websites for a living don't really understand how http requests are really processed, or how to read and use the chrome profiler and benchmarking tools.

        Claude isn't as good at optimisation work as someone who really knows what they're doing and goes deep on a problem. But I'm optimistic that it will help plug a capability gap in teams which don't have this sort of expertise on hand.

        (That said, the chance that people actually learn this stuff is going to also go down if people get used to outsourcing this work to claude.)

        • bluGill 8 hours ago
          The job of making a website is a large part something that should be given to people who have majors in art, English, and similar things. Well, there is some value in an engineer to build the whole site, but that's more of a framework task. The vast majority of the content is not an engineering thing to do. It is a job of people who are experts in language, art, journalism, those types of things that have very little to do with engineering.

          In short, if most of the people working on your website could do those things you named, you have a problem that you have hired the wrong people for the job.

      • latentsea 8 hours ago
        >The amount of complexity added for the gains is depressing.

        This has kinda been my experience too. I find sometimes it works and gets OK optimizations, and sometimes it "optimizes" things but in a really wrong way. It doesn't apply 'taste' to the optimization process to know what's appropriate and what's not.

        The way this will play out is people without the expertise will wind up using it to 'optimize' things and the next poor bastard is left to deal with the fallout.

      • jgtrosh 5 hours ago
        It shows the author have a deep understanding of the fact that premature optimisation is the root of all success
      • adamddev1 13 hours ago
        When people talk about AI being able to handle everything I keep wondering, have these people built anything complex, novel or serious? Just because people can see a website or a simple app improved, does that mean that all code can be handled by LLMs? It's like people are totally forgetting a whole category of careful, well-thought out programming for the critical parts.
        • whatisthiseven 13 hours ago
          Yes, what you are seeing are amateur developers that barely understand the tools they are using either giving LLMs poor instructions, or totally taking whatever it says at face value, then not bothering to put in further effort.

          If they just kept prompting it, or maybe used a different thinking level, it could have identified and solved this problem. Sometimes an engineer would look at a system and say "the current approach isn't delivering the desired engineering requirements. Maybe we need to rethink".

          Either engineer or LLM could take that sentence and run with it. OP of the article clearly can't do either.

          • theolivenbaum 13 hours ago
            You're missing the point where: in complex systems, sometimes optimizing code is both a high effort undertaking, and can totally not pay off. Having done hundreds of such exercises on our software over the years, it's liberating to have an idea of how to make something faster, being able to validate it without the fear of having to throw it all in the trash if it fails after days of work. What is still important is being able to provide proper guidance - we even built new tools to allow an AI agent to analyze memory usage in more depth, and instructions on how to benchmark in cloud environments where shared CPU usage and VM reallocation happen all the time and confuses the AI all the time with measurements
            • kanzure 9 hours ago
              Also, not every task is worthy of applying (human) galaxy brain consideration. Sometimes a task is satisfied by cheap + dumb work.
          • crooked-v 11 hours ago
            > the current approach isn't delivering the desired engineering requirements. Maybe we need to rethink

            I'm reminded of the times I've tried to let Claude fix some well-documented bug in the background, and it ends up burning 4 million tokens and 20 self-review cycles re-writing the same set of code a dozen times with ever more complex unit test mocks / overcomplicated regexes / giant comments restating the same thing the code does, when the actual fix turned out to be "change three to ten lines to do something in a slightly different way that avoids the problem entirely".

    • crooked-v 12 hours ago
      Even Next.js, for all that people are unhappy about its quirks, complexity, and random undocumented behavior and bugs (God help you if you ever want to try and actually use parallel routes as documented), can do all of this SPA stuff out of the box. Use `<Suspense>` appropriately on data loading, and everything static (including e.g. purely input-output components like editors) will load once and be re-used forever, and your Suspense-wrapped items will show a loading placeholder and only re-load when you intentionally re-trigger the data loading (which you can push down all the way to the level of individual buttons if you want).
  • minimaxir 15 hours ago
    This writeup legit coincidentally matches the asking-agents-to-make-code-faster-but-with-constraints-to-stop-agents-from-breaking-things writeup I posted on Monday: https://news.ycombinator.com/item?id=49803085

    Front-end UI optimization is slightly trickier than optimizing strict algorithms, but I found that prompts to the agents to build tooling to track visual regressions are more than sufficient. The main issue (at least with GPT models) is that you have to be very explicit about the use of padding/margins/negative space.

    That said, for my front end projects from scratch, I'm staying away from front-end JS frameworks and seeing how far and fast I can get with just HTML/CSS/vanilla JS shenanigans now that agents can wield them effectively.

    • fy20 8 hours ago
      > with just HTML/CSS/vanilla JS

      I was building web apps like this until 2017 when I entered the React + Typescript world. For B2B you can get pretty far, rendering HTML on the server is fast! I was using Rails on the backend, so templates, partials, shared chunks, made it easy to manage and have a consistent UI without repeating too much.

      The hard part is when you then need to build an infinite scrollable table, that has bulk select, and in-placs updating of columns. Ok maybe that's a bit too extreme the other way, but when you get to the point where it's easier to build full-on frontend components, you basically have to use a JavaScript framework for your entire UI. And they are basically all or nothing.

      Last time I checked (a good few years ago; I gave up and accepted un-optimized frontends as the rule) there wasn't really a good way to do progressive ehancement like the above: the page rendered on the server as HTML, and some components then become fully frontend managed. And no, frameworks like Stimulus and HTMX don't really solve it for me, I want something declarative.

      I'm a bit pissed off with DHH, that he went so far in the anti-Javascript direction, as IMO that was one of the big factors in Rails loosing it's limelight status. For backend I still haven't found anything as easy and fun to work with.

  • hungryhobbit 14 hours ago
    How about you make Opus 5.5 actually work?

    I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

    When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

    That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

    A model like that should never have gotten out of QA, let alone been released.

    • bitpush 14 hours ago
      I understand the frustration but shows a lack of critical thinking. Esp when you start with 'How about ..'.

      This blogpost is about frontend performance. It'll be akin to you commenting on a swift blogpost saying 'How about Airpods noise cancellation'. Sure both are Apple, but they are wildly different teams.

      • s3p 8 hours ago
        Through critical thinking I believe this would actually be akin to them commenting that AirPods firmware, written in swift, is bad because of swift limitations.

        In both instances, it's a side point that is actually tangentially related to the first. Not completely unrelated as you are implying

      • hungryhobbit 14 hours ago
        This is a techie discussion forum.

        In such a forum, it seems to me like it's fair game to point out that the company patting itself on the back about how great they are at programming (as evidenced in the article above about their 3x speed improvement) ...

        ... can't even make their latest model handle basic English without refusing to work.

        • jfidjcjwjcjwjd 13 hours ago
          Not that I’m defending Anthropic but OP is right, you’re complaining about the taste of a pear on a blog post about roses with the excuse that you’re in a biology forum. Sure, they both come from the same family, and sure, they both fall under the purview of biology, but they’re not the same.
          • s3p 8 hours ago
            I struggle to understand this argument. It is very likely that the user who is complaining about Opus 5.5 would be unable to do the very tasks anthropic is mentioning in their blog post, because of the refusals.

            In a blog post about Claude, i find it strange you get upset when people talk about Claude

    • dolmen 13 hours ago
      This issue is mentioned on the Opus 5.5 post [1] from Anthropic (no idea if it has been added after your rant):

        > Don’t ask it to show its reasoning in the reply
        >
        > What to do. Remove requests to reproduce its internal reasoning in the reply from your prompts and instructions.
        >
        > Why it matters on Opus 5.5. A request to reproduce its internal reasoning in the reply can be declined. It’s one of the flag categories.
        >
        > How. Ask Claude for what you need instead, for example, “Explain why you chose this approach in three sentences.”
      
      [1]: https://claude.dev/blog/getting-the-most-out-of-opus-5-5/
      • ronsor 11 hours ago
        Anthropic is trying so hard to "crack down" on distillation that they're ruining their own product. I do not know why.
        • PunchyHamster 9 hours ago
          AI companies are seeing the endgame, more and more tasks can be done just fine by cheaper model, and there will be not enough money to be made to support model that is only needed for top 1% of the tasks.

          And the goal is to slow competition down enough before IPO

      • crooked-v 11 hours ago
        That's extremely stupid, but also it's a completely different thing from literally just having the word "reasoning" in the text.
        • rendaw 6 hours ago
          If only we had some sort of artificial intelligence-like system that could decide not only based on the word but also surrounding context, and maybe even in cases when that word is not specifically used.
    • vikramkr 14 hours ago
      Probably it thinks you're doing some sort of system prompt exfiltration/distillation attack. Also what even is the workflow you're trying to have it do? It's doing code review but you're having it read some other AI models prompt/session history? Are you doing code review or like session history retrospectives?
    • railgunmerlin 14 hours ago
      seems a bit weird to complain about the model issues in a post about the harness/sites?
    • post-it 14 hours ago
      > When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

      Did it explain it did it hallucinate?

      • hungryhobbit 14 hours ago
        This happened at the classifier level, there was no "Claude thought X about it (hallucinating or otherwise)": this was a glorified regex deciding Claude couldn't work on a prompt (a code review prep) because it contained a string ("reasoning") it didn't like.

        It's more or less the same mistake we've seen Anthropic make repeatedly with it's brain-dead regex-based Fable/Mythos gates.

    • frumplestlatz 14 hours ago
      I’ve had the same thing occur five or six times over the past week; they seem to be attempting to prevent anything resembling chain of thought extraction.

      Every single time it triggered, it was due to a prompt written by their own model in a dynamic workflow. The self-serving nanny oversight has to go.

      The fact that they label model distillation as an “attack” is genuinely hilarious after they “distilled“ their models from all of our work, and continue to do so.

      I believe AI is here to stay and an incredibly powerful tool, but these companies, and especially Dario and Altman, are the very last people I want to see in charge of it.

      • copperx 14 hours ago
        > “distilled“ their models from all of our work

        They distilled all digitized human knowledge and artifacts and they're now complaining about someone copying their outputs saying it's a "national security concern."

        I'm not sure about how to classify that. Hilarious? Pathetic? Sad? Hypocritical? Hyperdramatic? All of the above?

    • Marciplan 14 hours ago
      [flagged]
      • hungryhobbit 14 hours ago
        Again, I used a slightly older model in the same series, Opus 4.6. It read the same prompt without any problem whatsoever. Also, a (non-Opus 5.5) Claude wrote the prompt in the first place.

        Opus 5.5 literally refused to work OR EVEN TELL ME WHAT I'D "SAID" when it read that prompt.

        Nothing to do with skill or the user at all: same exact prompt, three different models ... two worked, one didn't.

      • cyanydeez 14 hours ago
        the skill issue is "having to use a cloud model to do work of any value"

        might as well offer your life to a king to work in their fields.

  • pllbnk 14 hours ago
    > $500k engineer: [X] feels slow. Make it faster.

    > Claude: On it... Done.

    > $500k: Can you make it faster still?

    > Claude: On it...

  • simonw 13 hours ago
    I visited https://claude.ai/ over a mobile tethered connection from my laptop the other day and was pleasantly surprised at how quickly it loaded.

    (That said, I just had a look in Firefox and it loads 20.78 MB of JavaScript (6.84 MB compressed) so I expect they could make it a bunch lighter if they kept trying.)

    • mwcampbell 4 hours ago
      I doubt that they could iterate from where they are to an optimal solution though. What would have gotten them there more reliably is a development culture that would have found it disgusting to even consider shipping that much JS in the first place. And I don't buy the "then they never would have shipped at all" argument in this case, because, unlike the case of shipping cross-platform desktop apps before Electron, we did ship plenty of web apps, including interactive chat-shaped apps, before it became easy to unthinkingly ship a JS bundle that big.
    • brazukadev 2 hours ago
      > so I expect they could make it a bunch lighter if they kept trying.

      What is preventing a 1T dollars company from "keep trying"? What did they "kept trying" that got them to ship a webapp with 21mb of JavaScript?

      This is a failure of their software engineering culture. Keep trying the same thing will keep doubling their JS, not halving it. That's the effect of incompetence + LLM reliance.

      I wouldn't expect any better from a company that uses nextjs and React for their cli/TUI.

  • hmokiguess 13 hours ago
    You removed the load-bearing seams didn't you
  • whythismatters 14 hours ago
    The juvenile nonchalance with which some Anthropic employees seem to be talking to their AI (wacky, sick, cook, ...) is truly bizarre.
    • ieie3366 14 hours ago
      That’s just how people in their 20s casually communicate now. It’s mostly from tiktok
    • tantalor 12 hours ago
      They've never had a job at a normal company
    • ipdashc 10 hours ago
      It's a chatbot, not their boss. Who cares?

      If anything I'm glad to see them having fun with it instead of further convincing themselves that it's God.

    • unrented7977 12 hours ago
      That's literally just how kids talk these days. Are you really that unaware of modern slang?
    • queenkjuul 9 hours ago
      I'm pretty nonchalant with coworkers to begin with, when it comes to LLMs i sometimes am curious how well it can handle slang, and other times, since it's a program, am extremely rude. Their whole purpose is to handle natural language after all.
    • kylecazar 13 hours ago
      "You know what we want. Let’s go"
  • altern8 14 hours ago
    I used Opus 5.5 today for the first time hoping the writing would be more bearable and it SUCKS.

    Why can't they fix that

    • tarr11 14 hours ago
      Opus 5.5 writing is much more concise than 5.0
      • m00x 14 hours ago
        but Opus 5 is absolutely terrible. I'm still on 4.8 when I use Claude. I still think 5.6 Sol is the best available model outside of Astra/Fable.
        • copperx 12 hours ago
          Is Sol 6 terrible too?
      • altern8 14 hours ago
        MAYBE, but the bar is so low... It's not that it's good in any way
  • vikramkr 14 hours ago
    Less a post about performance and more about their Claude tag product. I guess it makes sense that there's not actually a ton of technical detail being moved into given probably opus is the only one that knows what all the dragons were lol - but cool workflow I guess
  • Tepix 2 hours ago
    What's absent from this report is the cost. From reading it, I don't think it was cheap.
  • felineflock 12 hours ago
    Claudehart's law says that when a measure is picked as an indicator, then Claude will start to game it.
  • tombot 12 hours ago
    Perhaps they can hill climb a native app instead of the electron and react native nonsense we have today
  • bouke 5 hours ago
    Interesting read. Could this be why recently support for Safari 26 has been going downhill? Especially Claude Artefacts where it either doesn't work, or says my browser isn't supported.
  • RomanKornev 14 hours ago
    The most important question is how much more unreadable the code became after all this "ratcheting the benchmark down". If you unroll a loop it will perform faster, but making changes to such unrolled code will be a mess. Will this make them ship slower overall? I'm sure at least half of it was just poorly written React code, but the other half?

    It's the same problem as overfitting in model training. If you're not measuring something it will get sacrificed.

    Or, perhaps the code quality literally doesn't matter anymore and we've reached "code quality escape velocity" where you can code as much slop as you want, the next generation of models will clean it up faster than the slop generates?

    • jackb4040 13 hours ago
      Came here to post this. AI rationalists love thinking about paperclip maximizers, but don't seem to care when it turns their own codebase into paperclips. Or to rephrase, turns their whole engineering org into meat proxies, slowing down engineering productivity in the long term because understanding is drained out of the staff and flushed down the drain every time they close a Claude Code tab.
  • pasteleft 8 hours ago
    Clearly the Antrophic don't even look at the single line of the code, so I'm curious why do they even need a Javascript (+ engine) for the Claude Code.

    They can just refactor it into C or Rust so they can ship a real binary instead of a script and a runtime.

  • khalic 14 hours ago
    Love it, the kids rediscover plain HTML and optimisation.
  • rancar2 13 hours ago
    Since I didn’t love the technical approach here (once one takes the humans out of the loop, there’s more ambitious things that can be done), I do appreciate the process. It’s not until the last sentence when process inspiration is revealed: “Special thanks to Boris Cherny for encouraging us to be more ambitious.”
  • mentalgear 11 hours ago
    So... this could optimize software so it runs again on older hardware, right ?
  • queenkjuul 9 hours ago
    It was seriously taking THREE SECONDS to render the chat prompt??? How do you possibly write an app that badly?
  • devin 13 hours ago
    I dare you to try and get heavy CPU cache-level performance optimizations you might see in tried and true HFT code written this way.
    • msephton 12 hours ago
      I'd be surprised if it wasn't possible
  • montroser 14 hours ago
    Okay, now fix the WYSIWYG markdown parsing in the chat input!

    Paste in a stack trace, then try to put it in a code block. Add a newline above, then add the opening triple backticks, then arrow down and add closing triple backticks at the bottom. Opposite congrats -- you have ended up with raw triple backticks at the top, plain text stack trace, and your cursor in a brand new code block at the bottom starting where you tried to close.

    Realize you want to go put code span backticks around some identifiers you wrote out earlier? Best make sure to insert them in the blessed left-to-right order, or else opposite congrats again -- you'll end up with a mix of raw backticks and code span treatment for the text between your identifiers.

    If Claude can discover novel CRISPR enzymes, surely it can make a rich text markdown editor, no?

    • AlexErrant 14 hours ago
      Here's a simpler bug:

      Claude Cowork still can't persistently reference a local dir, broken ever since they moved to their cloud project system, breaking many non-technical people's workflows.

  • geroge_kyaw 13 hours ago
    Sorry! I don't see that much difference. I would be more impress if you can make TTFT and ITL faster.
  • tecoholic 11 hours ago
    Great! Now can you also optimize the memory usage please?

    The blog post with all of the images, animations and embeds shows 167MB on my Firefox tab, while a plain claude.ai page is 297MB. For a text box with a few icons and some words, it a heck a lot of bloat.

    For context, my Slack tab with "Threads" open is 239MB. Yup, Slack of all things is using less memory than a zero-chat Claude.ai landing page.

  • chaordCAD 14 hours ago
    Great writeup really appreciate the detail on what actually worked vs. what didn't.
  • binlog 14 hours ago
    Step 1 - make a website that takes 3 seconds to load a blank page.

    Step 2 - bring it down to 1 second and pat yourself on the back.

    • xnx 14 hours ago
      Correct. A web page is fast by default. Everything that is added makes it slower.
      • joe_the_user 13 hours ago
        Sure, but some things increase the delay much than others. More static text, for example, is pretty fast. Elaborate JS, not so much.
    • airstrike 13 hours ago
      In Brazil they say you can make things better by first putting a goat in the living room. It's going to shit everywhere, chew, droll, and make a horrible mess.

      Take the goat out and your life will improve right away!

      • schoen 13 hours ago
        Also described as a traditional Yiddish folktale (a rabbi suggests responding to stress by moving a series of animals into the home, and then at the end they're removed and everything seems great by comparison!).
      • jfidjcjwjcjwjd 13 hours ago
        60+ yo Brazilian here and I have never, in my entire life, heard this saying before. Sounds something a Brazuca would come up with after the 17th round of Kaiser though. Pretty on brand.
      • Razengan 13 hours ago
        A life without a goat sounds horrible
  • bimguy 12 hours ago
    A few times now, I gave Claude all my IronPython scripts and asked it to make to see if it could improve the code. It proceeded to break nearly all of them.
  • dmvjs 9 hours ago
    just like its namesake Claude Shannon would say
  • j45 13 hours ago
    Faster is great, hopefully the quality remains.
    • queenkjuul 9 hours ago
      It took 3 seconds to load an empty prompt input: what quality?
  • owebmaster 14 hours ago
    With a simple trick: make it very slow first
  • dude250711 13 hours ago
    It can certainly measure your token spend rate.
  • wrxd 11 hours ago
    Startup time of Claude code is atrocious. I haven’t had a chance to have a look at what it is doing but it should be easy to measure
  • rvz 14 hours ago
    Let's try this again if you want an instant 10x speed up:

    Claude rewrite Claude Code from TypeScript into Rust. Make absolutely no mistakes.

    • cpursley 14 hours ago
      Not sure why you're getting downvoted, the Rust based tui's absolutely smoke Claude Code.
      • minimaxir 13 hours ago
        It's an overdone Reddit-esque comment that adds nothing to the discussion.
        • rvz 5 hours ago
          Nope. There's some truth in the contents of this article that anyone can read for themselves that any verifiable task that can be measured, Claude can optimize. Unironically, rewrites become more approachable and cheaper with a rewarding speed up if you know what you're doing; but not done all at once.

          So there's no need to get upset about that simple joke like the last time, and also complaining about comments being "Reddit-esque" which that breaks the HN guidelines [0].

          Lastly, you forgot to tell off another user who made the same joke today. [1]

          [0] https://news.ycombinator.com/newsguidelines.html

          [1] https://news.ycombinator.com/item?id=49820670

  • dccoolgai 8 hours ago
    Oh good, now Claude can commit the McNamara fallacy instead of my boss!
  • syngrog66 10 hours ago
    once I can measure something I can make it faster

    if I hand this challenge off to Claude then I will lose the capability to do it myself, and will become dependent on Claude. that is bad

  • datadrivenangel 14 hours ago
    AI written slop. They need to upgrade to Opus 5.5 or switch to OpenAI for writing.
    • joe_the_user 12 hours ago
      My user experience was significantly worsened just by their referring to "user journeys". I mean, I definitely don't want to go on a journey with Claude or ChatGPT (admittedly my go-to). I want result quickly and be done.
  • techpression 13 hours ago
    ”With that approach, we merged more than three thousand changes…” Why are they writing this? That’s terrible marketing all around, it means they let it go so far, with so little care, that they needed 3000 changes to make it into just a mediocre website performance wise (sure, electron app too, but still).
  • yunicc 9 hours ago
    [flagged]
  • sgarland 12 hours ago
    [flagged]
  • robertclaus 14 hours ago
    [dead]
  • dolmen 14 hours ago
    [dead]
  • boogiewoogie23 13 hours ago
    [dead]
  • boogiewoogie23 13 hours ago
    [dead]
  • applfanboysbgon 14 hours ago
    tl;dr if you make absolute dogshit software that takes 4.38 seconds to stabilize its first paint you can make really nice headline claims by "optimizing" it later
    • sashank_1509 14 hours ago
      Unironically true, can someone question, what is Claude desktop app doing that needs 500k+ lines of code?

      Agents complicate something that should be much smaller and simpler and then agents speed it up adding more complexity. I suppose functionally you may say this is fine but aesthetically it is hideous!

      • winstonp 14 hours ago
        the codex desktop app bundles libreoffice's libraries. i wouldn't be surprised if claude does something similar.
        • sashank_1509 11 hours ago
          This 500k lines doesn’t count libreoffice source code etc afaict.
      • bigwheels 14 hours ago
        TempleOS had kernel, compiler, 2D and 3D graphics drivers + libs, and tons of games and applications and the entire thing weighed in at less than 100kloc.

        AI coding agents of today definitely bloat things up [egregiously].

        • phoghed 11 hours ago
          That’s pretty cool. What are you using temple os to do?
        • theshrike79 14 hours ago
          On the other hand applications on TempleOS are expected/assumed to be perfect and have no bugs. It makes things a lot simpler =)
      • applfanboysbgon 14 hours ago
        There is nothing even functionally fine about this. The me who has programmed for a 4mhz computer with 128kb of RAM is crying inside. It's a fucking trivial interface for writing and displaying text and sending HTTP requests. You could have that displayed ~instantly on an 80s home PC. Now we have home computers that can execute somewhere between billions and trillions of instructions per second and yet a task that should take <10ms takes 4500ms. Our industry has become an absolute embarrassment.
        • adamddev1 13 hours ago
          Thank you. The apps created like this are absolutely an embarrassment. But to me the bigger horror is if people are going to be convinced to develop the critical libraries and infrastructure in the same way. Then we'll have bloat and rot in the deeper levels and everything will become exponentially worse and unreliable. It makes me so sad that people can't see this.
        • Daishiman 14 hours ago
          The Claude CLI does a lot more....
          • applfanboysbgon 14 hours ago
            This article is about the graphical interface, not the CLI.
            • aesthesia 13 hours ago
              The Claude graphical interface does a lot more...
              • applfanboysbgon 8 hours ago
                Sure, sure. There is still absolutely nothing it does that warrants such a slow paint (including the "optimized", still >1000ms paint).
                • aesthesia 7 hours ago
                  Definitely, but hyperbole about feature parity on a 4 MHz CPU with 128k of RAM doesn't help. I have conversations with Claude that wouldn't fit in that amount of memory.
    • Daishiman 14 hours ago
      Most line-of-business software has a lot of optimization opportunities. Making software optimized takes up time that can be spent building features. The fact that you can just make things go fast without having to take time away from feature building is actually pretty awesome.
      • applfanboysbgon 14 hours ago
        Not taking 5 seconds to load text is a feature. 10x more valuable than whatever other shitty feature you're thinking of piling onto your monstrosity. Like, actually tangible valuable to users and consequently your business; Google has already done the studies at scale that demonstrate how every 100ms of delay has an observable impact on usage statistics and user retention.
  • joyeljohn3 14 hours ago
    [flagged]
  • guideaitools 14 hours ago
    [flagged]
  • sonar_un 14 hours ago
    This was a fantastic read. Lots of useful info in there for your own projects.