Jeeves. Reasoning improves Jev-like decision models

(github.com)

63 points | by nicowaltz 1 hour ago

12 comments

  • RamblingCTO 14 minutes ago
    Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.

    But funny that jev is getting its lunch eaten apparently in under two weeks?

    • pavlov 3 minutes ago
      It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
  • thm 14 minutes ago
    Ask Jeeves - Only took us 30 years to come full circle.
    • tmnstr85 5 minutes ago
      this was the comment i came here for
  • swader999 32 minutes ago
    Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
  • Naitik88 8 minutes ago
    what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
  • captainbland 11 minutes ago
    See if it can beat Jev's Pokémon benchmark
  • mxkuzn 14 minutes ago
    interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
  • woadwarrior01 36 minutes ago
    This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
  • zerop 55 minutes ago
    Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
  • phplovesong 24 minutes ago
    So "askjeeves" has been resurrected?
  • raverbashing 50 minutes ago
    Jeeves, that's a name I haven't heard in a long time...
    • gizajob 30 minutes ago
      Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended.
    • fishfasell 44 minutes ago
      If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia
      • grokkedit 36 minutes ago
        jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves.

        soon I'll make sure that my home assistant pod answers to "Hey jeeves"

    • kjs3 13 minutes ago
      We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door???
  • alienbaby 1 hour ago
    Just curious, where has this term 'noul' come from for yes/no ansers?

    /a bit more digging and..

    A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.

    I hate it :)

    • doginasuit 18 minutes ago
      I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.
    • k__ 28 minutes ago
      The whole "no hallucinations" premise is based on that.

      Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.

      • doginasuit 14 minutes ago
        That seems like the only possible way to eliminate hallucination, short of a model that is never wrong.
      • kjs3 11 minutes ago
        force the user to decide in the end

        And that's...bad?

      • rusk 16 minutes ago
        Wait til you hear about how digital circuits work at die level
    • keepitwiel 1 hour ago
      Bernoulli
    • user3939382 1 hour ago
      If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this
      • kevindamm 56 minutes ago
        Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.

        In an analog circuit, maybe.

  • hjun1052 1 hour ago
    [dead]