Using an open model feels surprisingly good
Posted by msaltz 6 hours ago
Comments
Comment by mbanerjeepalmer 2 hours ago
I know the author has a long history here and I'm sure the post is a genuine reflection. But how is posting "I used my own product and it felt really good" anything other than self promotion?
Comment by witx 54 minutes ago
Comment by sd9 48 minutes ago
It doesn’t mean the post doesn’t accurately reflect the author’s views - it probably does. And the conflict of interest is well disclosed in the brief post. In fact, it’s quite natural for somebody who likes open models to work somewhere associated with them.
So personally I don’t find any impropriety going on here - but yes, it’s an ad.
Comment by Alisaqqt 25 minutes ago
Comment by fauigerzigerk 16 minutes ago
Comment by armchairhacker 1 hour ago
Comment by amelius 41 minutes ago
Comment by Doch88 25 minutes ago
Comment by dan_gee 2 hours ago
Comment by wps 5 hours ago
The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with it. If you’re using it as an aid to traditional software dev, iterating on small, targeted functions, GLM works as good if not better than Claude. Anthropic expects low quality prompts. If you rubber duck GLM, you get absolutely pristine output in most cases.
Comment by Gigachad 5 hours ago
Comment by wps 5 hours ago
I think the issue is that no one is content with incremental progress. We all know one shots are mostly possible, so the age of the personal project is kind of over. There’s no motive to invest dozens of hours getting a working prototype when Claude can give you something right now. So you can’t make small changes until you have that codebase in place already. You’re forced to make sweeping changes if you use AI from the beginning. And it’s not like it matters, there’s no personal attachment to any one part of the code, it’s not even seen!
Comment by Gigachad 5 hours ago
Comment by hangonhn 3 hours ago
Comment by techpression 2 hours ago
Comment by majormajor 4 hours ago
IMO because they speed-ran it so much (raising orders of magnitudes more, pushing out new stuff orders of magnitudes faster) it will also lead to the rest of the world catching up faster and shrinking the pool of people internal to them who continually benefit compared to what actually went on with IBM. There's only so much upside when you join a company with hundreds-of-billions valuation.
I guess the other part of their business plan bet was "become literal god of AGI". Maybe it'll still happen!
Comment by basch 5 hours ago
Comment by cootsnuck 4 hours ago
Comment by safety1st 1 hour ago
Claude feels more like gambling, that's what it is. Built for the vibe coder.
Comment by ThePinion 2 hours ago
Comment by cpursley 3 hours ago
Comment by Gigachad 42 minutes ago
It doesn’t have any subscription or limits. You just load up api credits.
Comment by kansm 4 hours ago
Comment by lordnacho 20 minutes ago
Comment by drnick1 5 hours ago
Comment by majormajor 4 hours ago
But the situations where you need that are narrowing every release. The "Composer 1" era of low-end models is pretty far away now.
Comment by Aquitard 4 hours ago
This has lowered the barrier of entry for many non-SWEs. I'm primarily a data scientist myself in the environmental sector with limited front end development.
My partner proposed an idea to help manage her horses and over the course of several weeks we fleshed out an android app that would enable/assist her with horse care.
I spent a few days in planning mode pointing claude to the services we want to utilise and the functionality.
We now have a fully developed app that we are testing but found several features missing. For example, Claude implement X but without a level of verbosity it failed to implement the edit/deletion of X.
The app is highly tailored to her use case and I had already built the main parts as a POC but she had no user-friendly method to interact with the information she required. Claude made this possible in such a quick time frame that just seems insane. As someone time poor (as most horse people are), it would have taken me a year+ to produce something unpolished when compared to what Claude produced.
Comment by Turskarama 1 minute ago
So not fully developed then.
Comment by dwedge 1 hour ago
This is my experience with using AI to generate tailored apps too. The first 80% makes me really excited but then it either only kind of works, has bugs, or has missing features. Even when I do eventually get it working I feel dirty using it because I just know it's badly implemented. Of the dozens I've created I don't think I still use any of them. I have a few glue scripts still in use.
Comment by scorpioxy 3 hours ago
To me it is great that LLMs are allowing more people to use computers and software the way they were meant to. But the people that are using them this way, in my opinion, wouldn't have commissioned anybody to do it anyway. They'd just live with whatever process/pain they have. So jumping from "I can now produce a prototype in a weekend" to "software development is dead" has always felt strange to me. Just something I've been thinking about lately.
Comment by Aquitard 2 hours ago
I understand what you're saying. My opinion is I see agentic workflow similar to the star trek universe where they ask the computer questions and get a response while they continue with their work. But that doesn't mean not learning how to do things from first principles. This is what I tell juniors in my field when I pass jobs to them. They can use LLM but to ensure they know and understand what's going on and most do from their university degree. I think this is where we need to pivot towards when discussing LLM.
Comment by scorpioxy 1 hour ago
Comment by MangoCoffee 3 hours ago
Comment by bananaflag 4 hours ago
After open models will also be able to do that, I expect everyone to forget "how software development is done".
Comment by theanonymousone 3 hours ago
Comment by mock-possum 3 hours ago
Comment by dan_gee 2 hours ago
Comment by arjie 5 hours ago
The TTFT and tok/s are much higher on the small model so that makes it competitive for a bunch of things. It feels like what old Sonnet used to by the end of last year which is honestly damned good.
Comment by pimeys 5 hours ago
I can do things like take a photo of a doctor's note among add the appointment to my calendar, send a PDF to my archive tagged, OCR'd etc, search info from internet, look data from Google Maps.
I can even integrate this to home assistant and talk to my agent with my open source Alexa-like system.
Comment by arjie 5 hours ago
In my case, I was foolish enough to run it all on my hardware which is pretty damned fast but has a duty cycle of 5% and runs idle most of the time. It's definitely better done via API.
Tell me more about the Alexa-like system! I have mine at home on a custom OpenWakeWord model trained on the word 'Aurora'. And I, too, have it hooked up to Home Assistant. I have a bunch of Eufy E21 baby cameras mounted on the walls[1] so it has vision through the house. The vision model uses GPT-5.5 on the subscription because I haven't yet set up a Qwen or Gemma multimodal that can see.
With Frigate on my home server I can even watch for events like my daughter waking up! And at night the agent sends out the vacuum if we've cleared the floor of baby toys. I feel this close to the dream of sci-fi AI. Because I auto-forward a bunch of my email to it etc. it knows about what's going on with me.
My wife will sometimes ask the agent information about me etc. and it's way easier to get a fast response rather than waiting for me to see the message etc.[2]
0: https://news.ycombinator.com/item?id=47538158
1: https://wiki.roshangeorge.dev/w/Blog/2025-12-01/Grounding_Yo...
Comment by Gigachad 4 hours ago
I can't wait until good enough models can run on affordable hardware.
Comment by ilteris 3 hours ago
Comment by arjie 2 hours ago
EDIT: Oh, home assistant. I use Home Assistant https://www.home-assistant.io/
Comment by mayank 5 hours ago
It was truly remarkably easy to build and package it exactly to my whims, in my case as a single docker container with a process reaper that runs llama, my Go code, tts, chat harness, browser in xvfb, and even a mailer daemon. With Gemma, it even runs on a RPi 5.
What’s absolutely wild to me is that over a couple hours, I could probably have it import parts of Home Assistant for my devices directly, and do other wacky stuff in what is essentially software for one.
Comment by loufe 5 hours ago
What would be useful is a cost metric. I'm curious how much I'd be willing to spend as a premium to not have those companies piping my conversations directly to the NSA. Maybe only some conversations? Claude and OpenAI are heavily subsidized, by all accounts, so Kimi K3 on a private endpoint might end up costing more or less - that's what I want to know.
Comment by mayank 5 hours ago
Comment by aprilnya 4 hours ago
Comment by lukeschlather 4 hours ago
Comment by m_ke 5 hours ago
Now I can't wait for someone to distill K3 into a Qwen 3.6 27b or Poolside S 2.1 sized models for a proper fast local Composer 2.5 replacement.
Comment by sudo_cowsay 5 hours ago
Comment by alpineman 1 hour ago
Comment by roywiggins 5 hours ago
Comment by sudo_cowsay 4 hours ago
Comment by khimaros 5 hours ago
Comment by wxw 4 hours ago
Harness-aside, I get this feeling sometimes when I swap from a big frontier model to something more nimble like Composer.
I can get into a better thinking and q&a loop with fast models, similar to how I can flow through a file more easily with vim.
Comment by madhu_ghalame 2 hours ago
Comment by shepherdjerred 4 hours ago
Comment by danny_codes 3 hours ago
Comment by bjertoref 30 minutes ago
Comment by seaal 3 hours ago
Really just feels like endless FOMO with how fast the iteration cycle is for harness and model development.
Comment by shepherdjerred 3 hours ago
I don't think it's very good though. As an example I can use one CC instance to delegate to several to achieve complicated/open-ended goals.
That just isn't possible with other harnesses, and it's definitely not benchmarked.
Comment by shepherdjerred 3 hours ago
With that said I think Codex/Kimi code are all behind CC as well, so maybe it's a question of effort and not the model.
Comment by spicyusername 5 hours ago
If it sounds good it must be good. The code this new model writes is incredible! It told me so!
Comment by adithyassekhar 4 hours ago
Comment by ljlolel 3 hours ago
Comment by _ink_ 2 hours ago
Comment by Alien1Being 46 minutes ago
Comment by luciana1u 2 hours ago
Comment by jdw64 5 hours ago
For example, if you modify things at the level of small functions, open models seem to perform just as wel
Comment by paradox460 5 hours ago
Comment by jdw64 5 hours ago
Comment by verdverm 5 hours ago
Comment by irishcoffee 6 hours ago
Comment by loire280 5 hours ago
Comment by c7b 3 hours ago
Comment by irishcoffee 5 hours ago
Making it about the hardware costs is not the play, they’re not an actual barrier.
Comment by hbn 4 hours ago
I'm not coding with Claude cause I enjoy the inherent novelty of LLMs. I'm doing it cause it's enabling me to quickly solve problems without getting stuck on the hitches that have deterred me from bothering with dozens of side projects my entire life.
Comment by kelnos 1 hour ago
(FWIW, I have the same attitude/preference as you do. Until I can run SOTA models on my own hardware without taking out a second mortgage, I'll pay our AI overlords for the privilege.)
Comment by anuramat 4 hours ago
but why?
Comment by irishcoffee 4 hours ago
Comment by alpineman 54 minutes ago
Comment by eru 5 hours ago
Comment by brcmthrowaway 5 hours ago
Comment by irishcoffee 5 hours ago
Comment by nujabe 5 hours ago
Comment by PontifexCipher 5 hours ago
Comment by nujabe 5 hours ago
Edit: I see, had typo last word of sentence was meant to be “blog”
Comment by kelnos 1 hour ago
Also, if you don't like a submission, flag it and move on. Commenting that you don't like it doesn't add anything to the discussion.
Comment by PontifexCipher 4 hours ago
Comment by sroerick 5 hours ago
Comment by Gigachad 5 hours ago
People need to learn to just post the prompt rather than the LLM output which just fluffs the prompt.
Comment by aprilnya 4 hours ago
Comment by capestart 12 minutes ago
Comment by wayknow 1 hour ago