36 comments

  • hadlock 20 hours ago
    If you're looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: https://github.com/Hadlock/laya-duum
    • sunbum 17 hours ago
      Why is this more impressive?
      • hadlock 17 hours ago
        It's not just shooting at static monsters in an empty room, this will go after monsters, make decisions about prioritizing health, changing goals, and finish the level, etc. It's not unique, it's just more impressive than the demo they're using.
        • runeks 15 hours ago
          Their demo could also change goals, e.g. avoid monsters instead of shooting them, by changing some variable. I'm sure that could also be used to change goals.
          • mosselman 13 hours ago
            But it doesn’t. That is why that other person is claiming their demo is more impressive.

            What you are saying is like me losing 13kg is as impressive as having the potential to do so if only I would stop eating ice cream and too many snacks every day.

            • trencedamp 11 hours ago
              You're confusing demo with technology.

              The potential to lose weight, in this context, is the original post because there was potential but it wasn't shown

              This commenter is the person who lost 13kg.

              The technology is not different but the demonstration is more impressive

    • antman 16 hours ago
      - Laya receives semantic snapshots, not the framebuffer.

      What does tjis mean? Another generative model?

      • ovi256 13 hours ago
        Laya is the model this uses. Like Jev, it's blind - it can't take an image as an input. So how can it play Doom, or Atari games, or do other visual tasks, as people have shown it to do? They write a bit of app-specific adapter code that transforms the game state into a structured piece of data (here, the "semantic snapshot"). And that is what these models take as input
    • prodigycorp 15 hours ago
      laya sucks. it's a dead end and for some treat it as a peer of jev or even the qwen jevs. base model doesnt have enough world knowledge to generalize.
      • earino 14 hours ago
        The base-model criticism is fair. Laya is weak zero-shot and doesn't have the world knowledge of a much larger pretrained model. Their own site puts the base models at around 0.35 on the typed-decisions benchmark, which is close to random.

        I don't think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark's train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That's a large gain, and it only shows up after you specialize the model.

        "Peer" is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn't compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.

        • prodigycorp 14 hours ago
          btw your comment was immediately marked as flagged/dead. Odd for a high quality comment. I vouched for it.

          My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you're finetuning, a bert base is not bad.

          • earino 14 hours ago
            Thanks for the vouching. I don't really understand what I did to anger the algorithms! But yeah, I think that that's the real sticking point. If the value prop of jev is "plug it in anywhere and it generally does smart things" you're absolutely right that Laya is Not A Thing. I've spent a bunch of time working on fine tuning models (i teach a class on it! https://earino.github.io/applied-deep-learning/ using the same ModernBERT stuff!) and so for me, I naturally went: "wait a second, so how good is Jev compared to fixing the problem domain."

            Anyways, thanks for the vouching!

            • prodigycorp 13 hours ago
              Great course, thanks for sharing, going through your slides and papers right now.
            • fwip 10 hours ago
              My guess is, it's because you sounded like an LLM in a few phrases there.
              • earino 9 hours ago
                I think that's a pretty good guess. I move around different countries a lot so I also assume I can look like a bot farm from multiple source IPs.

                However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I'm adjusting my speech patterns to its vernacular.

  • vektormemory 19 hours ago
    Can someone remove the extra LLM and just have an embedder do the classifier work?

    It's turning into pimp my llm...

    • dingody 16 hours ago
      I simply use an embedding model followed by a simple MLP, and that’s enough to solve many text classification problems.
      • 2gay 15 hours ago
        My little pony ?
        • fantispug 14 hours ago
          Multi layer perceptron. A tiny neural net.

          Boosted trees over embeddings work surprisingly well too.

    • nojs 19 hours ago
      Wait until you hear about support vector machines!
      • abhgh 19 hours ago
        Its funny - I was going to leave a similar comment - and I have, earlier, on a different thread. If people need fast classification, on a fairly scoped problem, it is very fruitful to start with an off-the-shelf embedding model like ModernBERT (which Laya uses) and stick a classifier in front - like a Support Vector Machine (SVM). For starters just tune the SVM, you don't even have to fine-tune the embedder - often it works very well, esp. given the compute needed. Plus you can get reliable confidence scores and generate explanations if you want them (using something like SHAP).
      • baobabKoodaa 15 hours ago
        If you have an easy problem that can be solved by a classifier from pre LLM era, then sure, go ahead. But we have LLMs now and we can use those to expensively and slowly classify harder problems using general purpose models, without needing to spend a huge amount of time fine tuning a model for the specific task. Jev offers to do the same fast and cheap.

        For clarity: no, a 0.8B model is not gonna do that.

    • last_health 14 hours ago
      [dead]
  • trebligdivad 1 day ago
    What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.
    • exogenousdata 22 hours ago
      The American stock market probably loses 20% of its value in a few days.
      • causal 9 hours ago
        Can you uhh elaborate on why?
      • drstewart 12 hours ago
        Can you share your short positions so we can see how confident you are in these predictions?
    • svachalek 22 hours ago
      I'd expect the vast majority at this point is coding. Classification is a thing but in my experience tends to run on light, cheap models, not the proprietary frontier ones.
      • RamblingCTO 16 hours ago
        part of agentic engineering is classification as well tho. reviewing/gating, what to read etc. don't need a full LLM
    • BowBun 21 hours ago
      We've stopped upgrading the models of our classification workflows for >1 year at this point, meaning they're running acceptably on early/mid-2025 models. That said I believe there is a long tail of non-production-ized users who throw this into their everyday LLM chats.
      • baobabKoodaa 15 hours ago
        In one of my client projects we've had to do forced upgrades of LLM models because of deprecations (I believe three or four of them over the course of 2 years). Each time our internal benchmarks have shown REDUCED performance after upgrading to "better" models.
  • AgentMasterRace 1 day ago
    I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
    • tbeseda 1 day ago
      For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

      I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

      Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

      • senko 1 day ago
        > The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

        You can already so that with classification models such as ModernBERT, at 0.4B.

        Jev's value is its zero shot performance without having to fine-tune.

        • beepbooptheory 23 hours ago
          I am sure I am missing something obvious here, but why is that valuable? Like, what kinds of projects are there where you need to classify stuff but are unable to make a bespoke model targeting the specific problem?
          • shaewest 23 hours ago
            For my org, it meant we could trial classifiers across various internal systems with little to no engineering effort. In one case we ended up building our own classifier instead of Jev, but in others we kept Jev because it was zero-effort for a great impact.
            • beepbooptheory 20 hours ago
              But like how often are you gonna do this in general? Why does dev time or effort really matter here when either way you are building something to just, you know, actually use going forward? Its not like one needs to build a new classifier everyday.
              • tyre 9 hours ago
                Because a lot of people employed as engineers can’t build classifiers. They’re able to glue libraries and tools together, build UIs, and write APIs, but don’t have the curiosity or creative problem solving to learn and master a new domain. Even when “master” is scoped to something like this.

                On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.

                I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.

                • beepbooptheory 6 hours ago
                  I guess this makes sense. I am no business guy, but if my product/company was focused on some sort of classification problem, my naive intuition would be to focus on hiring guys that can do it, rather than try to make the problem easier for them. But perhaps at the end of the day this is still cheaper? It's the same reason why we use Postgres instead of hire database experts to build something special?
                  • senko 2 hours ago
                    > if my product/company was focused on some sort of classification problem

                    It's more likely the product is focused on something else, but a classification model could come in handy...

          • senko 11 hours ago
            You need to have enough high quality data to train with, knowledge how to do it, and developer time.

            In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.

            I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.

            • davrosthedalek 8 hours ago
              Exactly. And that labeled data could be collected by just recording what they feed jev and what the decision is.
          • vickychijwani 17 hours ago
            Tons. For example, most web and mobile app developers won’t know where to start with making a bespoke model (and would likely have no interest in making one), but they will have lots of usecases for a classifier.
            • retinaros 12 hours ago
              its a 1 hour project on claude code on ur local machine. lol...
              • vickychijwani 8 hours ago
                Only for folks who already know what they’re doing.

                What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.

                It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.

          • jmalicki 22 hours ago
            > unable to make a bespoke model targeting the specific problem?

            There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.

            Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.

            For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.

            • benterix 16 hours ago
              The biggest disadvantage of Jev is that it's a proprietary product and you need to send them your data. A bespoke solution makes much more sense in many scenarios.
          • xigoi 15 hours ago
            Anything where you don’t have a decent amount of training data.
      • nico 1 day ago
        Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

        I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

        The classifiers also run in <1ms, so they can be very fast and precise at the same time

        But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

        For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data

        • derefr 1 day ago
          What do you use to determine that a particular task in a heterogeneous pile of tasks requires reasoning? The logistic classifier itself is too dumb to recognize the details of the problem that make it reasoning-sensitive (IIRC recognizing the “fiddliness” of a given problem requires a recognizer at least as complex as the problem itself.) And if you’re using the lightweight LLM for that, then you may as well skip the classifier and just use the LLM all the time, since that eval step is already going to be dominating your response time anyway.

          My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.

          • nico 23 hours ago
            I’m in the process of piecing together the different task/dataset-specific classifiers

            Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach

            For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time

        • rahimnathwani 1 day ago
          "Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification"

          Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?

          If so, I'm curious whether you compared that approach (A) with:

          B) Jev only, with a single output.

          C) Jev with multiple outputs fed into a logistic classifier.

          Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)

          • nico 1 day ago
            You are correct, I didn’t train the embeddings model

            Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

            The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller

            I haven’t compared different ways of sending requests to Jev

            The data to train the classifiers comes from the datasets used to test them (not from Jev)

        • jmalicki 22 hours ago
          I have not personally reviewed the benchmarks, but recall seeing a post where a Linear SVM with bag-of-words features outperformed Jev on many simple NLP classification tasks, like the ones you cite. You don't even need embeddings!
          • dcl 18 hours ago
            I've used LLM's to classify things for numerous projects and it's always really hard to beat linear classifier/decision tree over simple embeddings or even the sklearn HashingVectorizer. However, you DO need to trust your validated data, ground truths for this to work - but you should have these anyway to validate a Jev or similar solution.
      • clhodapp 1 day ago
        Needing fine-tuning for the use-case completely changes the product category
      • baobabKoodaa 15 hours ago
        You don't seem to understand what the point of Jev is when you say "you can fine tune it for your use case". Building your own classifiers for your specific business problems is the type of work we all used to do back in 2016 or so. It costs very much. Jev is a cheap and fast general purpose classifier.
      • gjs278 23 hours ago
        [dead]
    • Oras 1 day ago
      My use case is simple classification for job ads. Things like, industry, work settings (remote, hybrid, onsite) and job type (full time, part time .. etc).

      I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff

      I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.

      I suppose with larger model, this could be useful, but would require more ram and will be slower.

    • sharih 10 hours ago
      most of these dev clones need fine tuning on your use case, might as well fine tune modernBERT then. Jev generalizes well while being fast and cheap
    • zergrush 21 hours ago
      i've tried all the "open source" me too Jevs

      they all suck

  • olwmc 20 hours ago
    Sorry, do we have actual clear implementation details for Jev? I keep seeing these "recreations" or "Do Jev at home" but do we have access to their architecture? I haven't even used the product, I just find it strange.
    • Zetaphor 20 hours ago
      What they did was immediately obvious and trivial to replicate
      • prodigycorp 19 hours ago
        They make it clear that their edge is in their synthetic data. Train your own jevs miss the point. And if you have that much data, you didnt need jev or these replacements in the first place.

        HN's desire to pretend the data pipeline doesn't exist or isnt meaningful is silly.

        • Zetaphor 18 hours ago
          I'd argue that a tiny fine-tuned Jev on my dataset that I can run for practically free is far more valuable if the idea is to make this a dependency in real work
          • prodigycorp 18 hours ago
            People have been making their own calibrated classifiers and decision makers for a very long time. Mine outperform jev. But that's not the point of jev.
            • dcl 18 hours ago
              What is...?
              • kroaton 17 hours ago
                Being ZERO-SHOT. No fine tuning or re-training needed. Why are people so dense over this model?
              • ad1ttya 17 hours ago
                i guess it's not needing to finetune for each specific use case
                • jmpeax 17 hours ago
                  Then the probabilities are not calibrated. Jev is basically the LLM version of the "is-even" library.
                • dcl 17 hours ago
                  That is indeed reasonable. But you still need a bunch of trusted data to validate the accuracy of Jev.
                  • prodigycorp 17 hours ago
                    Yes. The difference is you can immediately benchmark and iterate via prompting without having to retrain. It's such a time saver.
                    • dcl 17 hours ago
                      Is it though? I have found it much easier to fit/diagnose/improve simple classifiers on embeddings/hash vectorized features/etc than iterate on prompts, especially true when using the more modern LLM's (like gpt 5.6 Luna) where you can't even set the temperature to get any sense of determinism.

                      I do note though, that is infinitely easier to 'deploy' a Jev/LLM based solution than a data/model pipeline.

                      • prodigycorp 16 hours ago
                        The comparison i'm making is against a training of encoder/small decoder model so I think we agree on most things. I dont think jev replaces the benefit of embeddings nor think they are mutually exclusive. All part of a handy utility belt.
  • phlipski 22 hours ago
    For those of a certain age - the fact that Askjev.com is still available astounds me.
    • taspeotis 22 hours ago
      I can’t believe ask.com threw in the towel in 2026. Like ride that LLM hype - how much better could the SEO for ask.com be? Do a pivot!
    • ozozozd 21 hours ago
      I am astounded you thought about this but didn’t buy it!
    • NetOpWibby 20 hours ago
      Blows my mind that Jeeves never came back as an AI model. Even if it's poor quality, it'd still be better than Google.
    • adefa 22 hours ago
      not anymore
  • ioma8 12 hours ago
    I made super quick file sorter with it.. Works like a charm: https://github.com/ioma8/jeff-sort
  • adrithmetiqa 1 day ago
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    • jubilanti 22 hours ago
      Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now, and open weight models that give you logprobs can give you distributions per key?

      Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?

      edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.

      • fzysingularity 22 hours ago
        FWIW you’re absolutely correct on how most developers are unaware of this as they’re mostly operating on the API level and not aware of server side (vLLM/SGLang) capabilities.

        The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.

        • sroussey 19 hours ago
          Classifiers (instead of LLMs) return results with confidence scores.

          Name Entity Recognition (NER) is one example.

          So many of them... https://huggingface.co/models?language=ner&sort=trending

          Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

          The nice thing about Jev is that people started taking about models that are not LLM text streams again.

          • anvuong 18 hours ago
            > Classifiers (instead of LLMs) return results with confidence scores.

            This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it's always stressed not to treat these sigmoid'ed or softmax'ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.

        • tclancy 22 hours ago
          Can you link to some info on this? I am just starting to poke around in the space and would love a leg up.
          • phoghed 22 hours ago
            This is an example implementation that was making the rounds recently. If you point whatever model you prefer at it, it’ll do a good job of explaining it.

            No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.

            https://github.com/Mushroom-Systems/lichen

          • fzysingularity 22 hours ago
            I don’t know of one definitive resource but “constrained decoding”, regex/grammar-based decoding, json-schema decoding will give you a bunch of hits. Look up vLLM, SGLang and outlines’ implementations for more technical details.

            This looks pretty decent: https://www.aidancooper.co.uk/constrained-decoding/

      • fooker 22 hours ago
        > Like what am I missing?

        Orders of magnitude faster and cheaper answers.

        Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?

        • vengadanathan 20 hours ago
          No necessarily, we are getting similar result gemma 4 12b, predicting one output token for label. it is just prefill and with nvfp4 caliberated , we are getting 70-80 ms p90 for the classification workload we have.

          not sure what made you think you cant go cheaper than jev. it is being done for a long time.

          • fooker 8 hours ago
            Can you host files cheaper than Dropbox? Of course you can. :)
        • bitpush 21 hours ago
          Excellent analogy
      • Lerc 21 hours ago
        I remember reading a paper about a model that used a parser bound to the end of a LLM where it squashed all logprobs for outputs that were parsing errors. I have not seen this used in the manner I thought I would. I thought it would have been entirely possible for a model to construct it's own grammar for how it would prefer to respond and then opt to generate tokens that matched (with a metacode to turn it off obviously).

        That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.

      • sheepscreek 22 hours ago
        Many people yes, but it’s probably for the best. Without fine-tuning on such style the results would be unreliable. It wouldn’t be the probability of the outcome of whatever you intend, but just the next token probability - which could alter if you add a space or punctuation to the prompt. Very flaky.
      • peab 21 hours ago
        Yeah, I'm with you.

        Nowadays any LLM and any harness you use will just do this for you.

        But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.

      • jrop 20 hours ago
        Yeah I though llama.cpp had this during the very early days, if my memory serves me correctly.
      • nextaccountic 22 hours ago
        Pricing

        The same reason I want diffusion language models to be mainstream

      • est 22 hours ago
        You don't even need a constrained decoding.

        As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.

      • sampullman 22 hours ago
        Price and speed are the difference. But deepseek flash is usually fast and cheap enough, so the use cases are somewhat limited.
      • ed_mercer 22 hours ago
        jev is optimized for it, a standard LLM isn't and it also costlier and slower.
        • fzysingularity 21 hours ago
          When you say Jev is optimized for it, I get that there’s no need for unnecessary decodes in an autoregressive fashion.

          But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.

          We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.

    • tbeseda 1 day ago
      If I had to guess, it's already built and is just waiting on Product's/Marketing's desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don't want to adopt the same verbiage+acronyms - but also can't be seen to be just sherlocking features.
      • seizethecheese 1 day ago
        I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.
        • bigyabai 1 day ago
          I think Nvidia has demonstrated that shipping first is a multi-trillion dollar opportunity if you don't shy away from a challenge.
          • lucideer 1 day ago
            I think it depends on the model: Nvidia ship products with APIs, apple/jev/etc. ship end user products. The former is much more subject to lock-in, increasing the value of early market adoption because there's a 3P Nvidia ecosystem sprung up in response. Apple/jev/etc. do have APIs & corresponding 3P ecosystems but those are usually a smaller component of market capture than direct product end users, so the space ends up more competitive.
            • 8note 21 hours ago
              jev ships an api, does it not?
              • lucideer 8 hours ago
                yup. I said that in my comment...
          • AndrewKemendo 1 day ago
            NVIDIA literally hired all of the 3DFX team and patents after they already proved the graphics accelleration hardware market was massive with the Voodoo card line

            In fact NVIDIA wasn't even a close competitor to 3DFX in the graphics card game at that point

            • bigyabai 23 hours ago
              It panned out great. 3DFX filed for bankruptcy less than 18 months later, and Nvidia could pivot from designing raster chips to considering CUDA's architecture.

              It's not like 3DFX was the first GPU vendor. Nvidia saw the opportunity to be the first true GPGPU vendor, and they beat their competitors.

              • trollbridge 23 hours ago
                IBM shipped the first PC GPU (the Image Adapter/A, 1989) following on the first PC 2D accelerator (the 8514/A, 1987). There was zero benefit to being first. 3dfx had the first mass market, cheap GPU in 1996.

                ATI (AMD) in 1989 copied the unpatentable parts of the 8514/A, improved it, and went on to dominate 2D accelerators.

                nVidia’s first 3D card was a complete flop. They did not achieve success for years.

                Nvidia is in the right place at the right time.

                • Keyframe 21 hours ago
                  not to be _that_ guy, but S3 was the dominant force in 3D. ATI was nowhere to be seen with their Mach/Rage crap. It was an S3 era, and then seemingly out of nowhere 3dfx swept in with Voodoo (not seemingly though - 3dfx came out from imploding SGI). Only a bit later Nvidia, after it recovered from NV1 fiasco, brute forced relentlessly from Riva 128 to TNT to TNT2 to Geforce 256 which broke 3dfx (along with their lack of business sense) and had Nvidia bought them. ATI did their 2D schtick only during that timeline while S3 stumbled with Virge and ATI didn't come as a threat until Radeon; After they bought ArtX (which did Gamecube graphics) which turned into Radeon. S3 died off in that race with Savage3D and the only other major players were Matrox and 3dlabs. Matrox kind of found temp refuge in video segment, and 3dlabs infamously pushed for OpenGL 2 and survived for a bit on 3d workstations. The most surprising (IMO) was the downfall of E&S which kickstarted most of the things mentioned. E&S and Real3D are both a great story in themselves how first movers can become absolutely forgotten and obscure real quick (with Intel740 ended up in ATI).
                  • trollbridge 11 hours ago
                    S3’s stuff was no more “3D” than the Image Adapter/A was.

                    You should look at the IA/A - it had its own C like compiler, CPU, etc which did things reminiscent of a modern GPU or SIMD.

                • bigyabai 22 hours ago
                  We're talking about GPGPU products, not raster GPUs. I cleared that up pretty well in my last comment.
          • PunchyHamster 23 hours ago
            It's more due to ineptitude of competition. If AMD shipped second, but better product it would be another thing but it is still a bit of a mess of an ecosystem on AMD side
            • bigyabai 22 hours ago
              It's not AMD's responsibility to dethrone Nvidia any more than it is Apple's. AMD sells CDNA, but the momentum is with CUDA and Khronos can't get anyone to sit at the same table anymore.
              • hgoel 20 hours ago
                It isn't their responsibility, but it doesn't make sense to argue that being first got NVIDIA a multitrillion dollar market if the others aren't even trying to compete. There is no "first" if it's really just "only one even trying".

                The momentum is with CUDA because CUDA is the most broadly usable one. Especially with AI-driven optimization loops and similar APIs, competitors can more easily pick up momentum, if they'd actually try.

          • throwaway27448 1 day ago
            Chip manufacturing intrinsically comes with one hell of a moat. There's not much parallel in software.
            • slashdev 1 day ago
              That’s kind of funny because Nvidia’s biggest moat is arguably CUDA, the software ecosystem around their chips
              • angry_octet 1 day ago
                Historically there was a big patent moat in (Graphics) GPU design. This continues with CUDA, but obviously Intel and AMD could find ways to support eg PyTorch. What we don't know is how much effort that cost them, or why they decided they couldn't make a CUDA API compatible competitor.
                • trollbridge 23 hours ago
                  Intel could have beat the pants off Nvidia a long time ago with Arc if they’d bothered to ship usable drivers. But they refuse to, and simply can’t figure it out, so Arc cards remain cheap because they’re so #%£€ing hard to get working well, and everyone is nervous they’ll lay off the driver team again.
                  • api 23 hours ago
                    AMD has always had software problems too. Hardware companies often devalue software and suck at it.

                    If you can make them work Arc cards are a massive bargain. On raw compute the silicon is not bad.

              • mcmcmc 1 day ago
                CUDA is a lock-in moat, the infrastructure needed for chip manufacturing is a barrier-to-entry moat. Two different things.
                • angry_octet 1 day ago
                  There are many microarchitecture patents used in NVIDIA chips. I'm sure they have a team that rips apart AMD chips looking for infringement. The way CUDA works is tied to many GPU architecture decisions and it would be hard to decouple them efficiently. Obviously a huge effort was made to get PyTorch decoupled from CUDA.
                • AtlasBarfed 23 hours ago
                  If cuda is an API, and llms make apis effortless, then how big of a moat is cuda?
                  • latentsea 20 hours ago
                    It's becoming less of one. Previously I would have shied away from buying an AMD card because of CUDA, but with local LLMs getting good enough to be usable and frontier models becoming as good as they have, I bit the bullet and got an R9700 for local inference. Dealing with working around CUDA used to be more of a manual process, but when you can point an agent at it and get stuff working, it's dramatically less painful and scary than it used to be. Plus, at least in ComfyUI and local LLMs I'm finding support for AMD has gotten really good. Lately I've been witnessing a lot of people using agents to write custom kernels for RDNA4 and improving performance dramatically.
                  • bobmarleybiceps 23 hours ago
                    IMO, yes a lot of the "nvidia pays lots of people to make non-portable, tightly coupled backends to open source projects" is potentially going to be less of a moat?

                    (Though it could turn into "nvidia pay lots of people to use LLMs to make non-portable, tightly coupled backends to _even more_ open source projects")

                  • trollbridge 22 hours ago
                    It turns out the moat is “writing drivers that work”; Nvidia drivers simply work, and Intel’s are poor quality. So if I want stuff that works I need to buy Nvidia gear.
                  • robflynn 23 hours ago
                    I ran across this a few days ago: https://zluda.org/
                    • QuantumNomad_ 19 hours ago
                      All of the buttons and links on that page redirect to spam pages. Most of the times I clicked, it brings me to some site that wants to sell me a VPN.
                      • robflynn 8 hours ago
                        Oh, yikes, I should've checked those links before posting it here. That's certainly not a good look for them.

                        There was a github repo but I have not checked it.

                        edit I see, thats an unaffiliated site that latched onto that, my bad, here's the GH that I should've linked: https://github.com/vosen/ZLUDA

                • flyinglizard 22 hours ago
                  I’m totally guessing, but I can’t imagine CUDA has any significance at the frontier lab scale. The operational and capex costs are so massive that the convenience of the platform becomes a minuscule consideration.

                  It’s just that Nvidia’s stuff works, and available at scale, and includes the full stack with networking, cooling and such.

      • michaelrwolfe1 22 hours ago
        There is zero stigma to shipping second. If anything, the labs’ customers are probably begging for them to add these features natively so they don’t have to deal with the hassle of adding another provider to their stack.
    • wgd 23 hours ago
      Negative three years, give or take. Although recent Anthropic and OpenAI models no longer expose the capability. But for any open model you just tell it to respond with a single token "Y/N" and take the logit difference. If you want multiple distinct questions answered you just ask them independently and put the shared context first so it gets cached.

      OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.

      • Renaud 23 hours ago
        I think the main differentiator offered by Jev is not the ability to answer questions, most models can be coerced into that function if they don’t already have a dedicated pipeline for it, rather it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.
        • selcuka 23 hours ago
          > it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.

          It also returns confidence scores for all choices.

          Granted, they are not stable. They fluctuate even when you reorder choices, but it still counts as an additional feature.

          • wonnage 22 hours ago
            They’re useable if you’re trying to rank autosuggestions or something, not great at implying actual understanding
        • wgd 23 hours ago
          The evaluation is fast because it's all prefill computation with only a single token of inference. Ditto cost, you're paying 100% input costs and nearly zero output. There really isn't any architectural magic to Jev, it's just a straightforward application of normal LLM tech with some good marketing.

          I mean, Jev is also probably cheaper because it's a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.

          • psyphy2 23 hours ago
            Yes this is exactly my assumption as well. I think ppl forgot before agents it was expected slo to have a ttft in range of a few hundard ms, which is what jev is also achieving.

            Thats also why it feels weird they say they don't "charge for output tokens" since its literally generating a single (or at most very few tokens).

    • k__ 1 day ago
      No need, as that functionality can run locally no problem.
      • make3 1 day ago
        the appeal would be if they can deliver it at a much higher performance and similar speed, which is plausible
    • onlyrealcuzzo 1 day ago
      Probably at the frontier stage - you will only see it where Jev is better regardless of cost.

      For everyone else who is conscious of cost, you're already seeing this being built into harnesses.

      Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.

      If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.

      I would be astounded if Anthropic leads the way on a cost reduction.

    • bigyabai 1 day ago
      BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.
      • alanwreath 1 day ago
        This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?
        • yfontana 1 day ago
          Jev is a classifier. The big thing about it is that it has high accuracy on domains it wasn't fine-tuned for, like an LLM, but with speed and cost comparable to traditional classifiers.
        • fra 1 day ago
          BERT need to be fine tuned for your use case, Jev generalizes. It’s a pretty big difference!
          • latentsea 20 hours ago
            So what you're saying is it's artificial... general... intelligence? /s
        • _menelaus 22 hours ago
          The fact that its not narrow and stupid is what's different
        • make3 1 day ago
          FLAN-T5 generated text (Jev does not generate text), and BERT wasn't able to do tasks without fine-tuning.

          Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.

          Eg, the following query is put in the encoder model:

          {"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }

          The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.

          The UI stitches the first 5 floats with the choices and returns something like:

          {"rank": ["ant", "mouse", "cow", "truck", "tower"]}

          • vlovich123 23 hours ago
            How does it know that you’re asking for “rank” instead of something else if it’s not generating text?
            • jmalicki 22 hours ago
              By reading, not generating, text.
              • vlovich123 20 hours ago
                It has to output “rank” - that requires generation unless I’m mistaken
                • make3 20 hours ago
                  no, 255 numbers come out at once all the time, the order is determined by the inputs
                  • vlovich123 20 hours ago
                    That’s not what I’m asking about - the text prompted for rank and it output “rank” in the response object. How did it do that? Like if I’d asked it to group similar items, how would it know how to structure that output and know that the key should be “groupings”
                    • make3 7 hours ago
                      The output key is determined through code by the harness from the inputs, and the model generates 255 floats in order that follow that schema by reading the expected return type "rank" in the input. The harness then programmatically uses the floats of the 255 floats that are useful. The harness can assume that the correct float will be in the correct position as the model is trained to follow the schemas
                    • pests 16 hours ago
                      [dead]
          • aftbit 21 hours ago
            Is Jev's architecture public somewhere? I'd love to read more about it.
          • dannyw 1 day ago
            So it generates logits in a 255 token output space? ;)
            • make3 23 hours ago
              logits assumed some form of softmax or logistic, which may not be the case
        • prometheus1992 1 day ago
          nothing in utility. we used various bert variants to satisfy our usecases and still in uses. free and they run locally.
    • Ohentis 1 day ago
      I don't think there would be any utility for that. Anything jev can do, a frontier model can also do. Just not as quickly or as cheaply.
      • tbeseda 1 day ago
        I think these products (Jev and the inevitable offerings from Anthropic, OpenAI, etc) want to become more than end-user output machines. They'd benefit from being in the hotpath of other services. Not backgrounded generation but in-band, request-time work.

        1M x $0.50 == 1B x $0.0005

        • transitorykris 1 day ago
          To expand a bit for my current use cases. Inline routing of work to heavy task specific models, and prompt/context generation (user is asking something, what and how much should we prompt the expensive LLM with). Latency or time to first token does matter for some applications.
        • Ohentis 0 minutes ago
          [dead]
    • themitchelli 1 day ago
      But for those of us that prefer open source and self hosting, JEV alternative LAYA will beat anything the frontier models package up.
  • danbrooks 1 day ago
    This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.
  • m11a 16 hours ago
    It sounds like Jev is not a generative model. If true, I don’t understand these clones which fine-tune a generative model, which appears to not be what Jev is doing.
    • blueblazin 15 hours ago
      I think you may be confusing autoregressive and generative. But that can just mean it's a BERT-like encoder model that outputs a single vector for the given input tokens. To make it sound even more "novel" you can say it's a bidirectional encoder-only transformer model which modern LLMs are not and sell that as something revolutionary.
    • petesergeant 16 hours ago
      > It sounds like Jev is not a generative model

      Why does it sound like that to you?

  • jiwidi 12 hours ago
    I wonder, how different is this now from qwen trained rerankers like zeroentropy ?
  • joe-excom 16 hours ago
    I see you were turned away from submitting a "Show HN" too. I might have to do the same as you for my project.
  • pan_lid 19 hours ago
    30ms for a 0.8B decision model is crazy fast. Most of my inference pipelines struggle to hit that with smaller models.
    • anvuong 18 hours ago
      Is it though? I got sub 10ms for 200M params vision model way back in 2020 with TensorRT optimization running on RTX 2080Ti. 30ms for 0.8B classifier model doesn't sound special to me.
    • prodigycorp 19 hours ago
      how many ms is it for a 32k token prefill?
  • Kvarnek 20 hours ago
    Training 0.8B models at home with that latency is seriously impressive. What kind of hardware setup did you use for training?
  • imranq 21 hours ago
    Everyone saying you could replace Jev or decision type models with an LLM with bolted schema output constraining are missing the point completely. Its about extreme speed and cost effectiveness with high quality, neither of which you are going to get with LLMs even with these KV-cache tricks
    • computerex 19 hours ago
      There is no free lunch. You cannot get high quality on open domain and have it be fast.
  • torutofu 10 hours ago
    neat if the interesting bit is when the 0.8B call disagrees with a dumb heuristic on the same features, not the headline latency.
  • qtalen 22 hours ago
    Awesome, I was just looking for a decision-making model that can be deployed locally, and here you are. Thanks a lot!
  • k__ 1 day ago
    Von 1.2 had a better Doom score :D

    https://github.com/wfzyx/von

    • nico 1 day ago
      Oh wow, those Doom scores for Jeff are pretty terrible

      The Von numbers have led me on a rabbit whole of getting a classifier to play Doom

      I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)

      Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)

      It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU

      Edit: after looking at Jeff's numbers more in detail, the 6.5 kills number is not that bad, but it can definitely be better ;)

      • kridsdale1 1 day ago
        One day we’ll be using the same Kills/SOTA metric for models driving physical kill-bots.
  • folayii 20 hours ago
    83.1 vs 83.0 on your panel against 70 vs 94 in someone's actual use case is the whole story with zero-shot classification. Any sense of what the 0.8B does on a phone NPU instead of an M4 Max? That's what decides if it's shippable on device.
  • port11 13 hours ago
    Is this Jeff Vader, head of Catering?
  • velominati 1 day ago
    Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
    • odo1242 1 day ago
      Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)
      • Ohentis 1 day ago
        I feel like the most likely is that it largely works like how all the recent copycats work: take an existing llm, modify the decoder, do RL training. I think the main reason that jev works better is that they spent more time on that post training step.
        • odo1242 23 hours ago
          My guess as well
      • microtonal 18 hours ago
        But nobody runs the whole model again for each predicted token in each decode step, this is what KV caching (and prefix caching) is for, to keep the model running in O(n^2) overall (and O(n) per decoding step).
        • odo1242 7 hours ago
          Models are O(n^2) on input size even with KV caching - the KV cache only decreases the processing time by a constant factor.

          (Specifically, there are two O(n^2) steps in an attention layer and KV caching makes the first one O(n) with caching - but the overall big-O is still O(n^2) because KV caching doesn't affect the second step.)

          • microtonal 5 hours ago
            I think that I understand where you get the idea from, but you are mistaken. Check equation 1 of the attention paper:

                softmax(QK^T/sqrt(d_k))V
            
            So without KV-caching, if suppose you have N tokens, then you have

                Q: N x q_dim (leaving the batch size and n heads out for simplicity, they are constants for our purposes here)
                K^T: k_dim x N
            
            So

                QK^T: N x N <- this is the attention matrix
            
            The scaling and softmax are irrelevant here, since they are elementwise

            Then

                V: N x v_dim
                (QK^T)V: N x v_dim
            
            This gives us the normal quadratic complexity of prefills (of course, in an actual implementation an attention mask is used to ensure that tokens cannot attend to prior tokens and you may have things like ALiBi).

            Note that during decoding, the dimensions change. Since it is an autoregressive model, we do not need to recompute the values and keys of prior tokens, only of the token that we are currently decoding. Of course, the token that we are decoding still attends to all prior tokens and itself.

                Q: 1 x q_dim
                K^T: k_dim x (N+1)
                QK^T: 1 x (N+1)      <- Note that attention in this step is not quadratic anymore, since
                                        we only need to compute how the current token attends to prior
                                        tokens, the representations of prior tokens are frozen. K comes
                                        from the KV-cache.
                V: (N + 1) x v_dim   <- V comes from the KV-cache
                (QK^T)V: 1 x v_dim   <- Also not quadratic, the value is only computed for the current token.
            
            
            So attention during a decoding step is O(N), so when decoding N tokens, it is O(N^2) overall.

            The point-wise feed-forward layer does not matter, in decoding it only needs to be computed for the representation of the token that we are currently generating. We don't need the representations of the preceding tokens for the next layer, since we have already cached their keys and values for each layer.

            Disclaimer: I was one of the developers of a widely-used inference engine and implemented several of these optimizations.

      • k__ 1 day ago
        Then again, it's only a very small number and fixed set of tokens for the output.
    • hellohello2 23 hours ago
      The main difficulty with fast (low-latency) inference, is not actually computation but loading parameters from memory. The problem with generating 1 token at a time isn't that that its expensive computationally (it is, but so is training), but that you need to stream your entire model from memory for every single token (and also the KV cache but that's besides the point). So the strategy usually taken is to process big batches of user requests, letting you share the memory loads across users. This means individual answers aren't that fast, but you can do lots at once. This is why local inference isn't cost-effective; its because the model weren't designed for it in the first place.

      My hypothesis for Jev is that they simply generate many answers independently in parallel from your prompt, and then discard the duplicates (or train to avoid duplication with attention between the ). In that way the entire batch is 1 user's prompt.

    • mrbonner 22 hours ago
      My guess is that they use an encoder-only model as the foundation and then do RLDC. Why? 1) it doesn’t need text generation 2) limited context window (40k last time I check)

      Those are telltales of a Bert model.

  • ijustlovemath 1 day ago
    Isn't jev just a less nuanced classifier? What am I missing?
    • neuronexmachina 23 hours ago
      The zero-shot learning is what makes these different from a traditional classifier.
      • 0x457 23 hours ago
        Zero-shot just means you give it zero examples. Jev lets you add examples, so it's one-shot/few-shot depending on how many you provide. After playing with it, it seems like once you wonder off their guide examples domains - you have to provide examples to get anything useful from it.

        IMO you going to get better results by doing a small tune of a tiny model. Making training dataset for it with LLMs is easy, serving it is going to be cheaper.

  • drzhouq 1 day ago
    Anything like this in the VLM side? Classification on images...
  • kdaniel_03 16 hours ago
    Local models are getting better, its going to win
  • zeroCalories 23 hours ago
    I've always been more afraid of these types of models than LLMs. These are what enable mass surveillance at scale and autonomous real time combat drones. Now they are spreading and being optimized. Gg.
  • bilekas 1 day ago
    Can we get a price comparisson ?

    Edit: Running them for the masses.

    • buildwithdennis 1 day ago
      I assume the cost is whatever you run the model on. Jev is already dirty cheap, $0.42 per million input tokens and output is free and the input cost is covering tokenization + API utilization
      • bilekas 1 day ago
        Well, I'm no expert, but this runs on Qwen3.5 and Gemma 4, stripped down, but they're pretty pricey.
      • rahimnathwani 1 day ago
        $0.042 not $0.42
  • yesthisiswes 1 day ago
    My name jeff
    • esafak 22 hours ago
      No, he jefe, man.

      I've been holding it in since I saw the title. I was surprised it wasn't all over repo!

  • desireco42 22 hours ago
    I think Jev is still the king and while I really appreciate this and other similar projects, Jev gives assurance and value that is hard to beat. Well.. this works locally which is always best, even if slower.
  • lin7c 9 hours ago
    [flagged]
  • firelex 1 day ago
    Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe).

    When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.

    The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.

    The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:

    - Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.

    - Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).

    - Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.

    Videos of every run are linked in the README.

    Lessons learned:

    - System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.

    - A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.

    - Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.

    - Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.

    - Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.

    - Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.

    • ironqcold 1 day ago
      Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?
    • thomasikzelf 1 day ago
      Great project! How does this compare to asking qwen to reply with 1 token in terms of speed?
    • sorenjan 1 day ago
      That's some local hardware.
      • e12e 1 day ago
        Was going to make the same observation. Cool to run locally - but renting seems the saner choice?

        Would love to see what this run would cost from something like Verda. Just out of curiosity. I'm not going to be installing any sparks at home anytime soon.

  • pwython 20 hours ago
    [flagged]
  • armanj 1 day ago
    [dead]
  • TechnoBabble199 22 hours ago
    [dead]
  • zane_shu 1 day ago
    [dead]
  • basil_io 19 hours ago
    [dead]
  • badatnames 1 day ago
    [flagged]
    • tintor 1 day ago
      First commit of six commits was 4 hours ago.
    • behnamoh 1 day ago
      I'm curious: in the age of AI, where you can literally tell your clankers to work on something, why do you even care about the longevity of open source projects? There are several projects that are abandoned, and I have resurrected them for my own use cases without any issues.
      • remywang 1 day ago
        By the same reasoning, why should I care about the output of someone else’s clanker if I can just get it from my own clanker?

        But to answer your question, I do still care about software being maintained by someone who can make design decisions instead of just yielding control to bots who tend to produce mediocre designs.

        • rukuu001 1 day ago
          These are all good questions.

          Perhaps we can think of projects like these the same way we’d think of a good tweet - not valuable in itself but valuable because of the ideas or conversations it generates.

        • devin 1 day ago
          Maybe because they know more about the domain than you do, and it will take you a lot more back-and-forth with your clanker than it would for them with their clanker.
          • stymaar 17 hours ago
            And maybe they don't, and that's the problem with one-day-old projects: we know for sure that the programmer didn't have time to learn anything from that particular project and we now have to trust them that they know what they are doing. Unfortunately this is wrong in 90+% of the time, so that's kind of a gamble to expect one particular project to tick this box.