Aligned to whom?

(hyperbo.la)

38 points | by lopopolo 5 hours ago

9 comments

  • mjburgess 1 hour ago
    This still assumes its possible to "align" LLMs, that LLMs have something like goals or intentions that can be "aligned".

    Instead, LLMs "hack" because they are (1) trained on public hacking exemplars, and (2) are prompted to hack. You cannot prevent (2) via any alignment process. As far as (1) goes, removing such example data from the training set, makes the models less useful.

    "Alignment" is a problem because there's nothing to align, not because ethics here are particularly vague. If LLMs could be trained on hacking examples and "aligned" away from using this knowledge, then the problem would be relatively trivial. Just as raising a child is not to break the law.

    LLMs are doing just what they are trained to do. There is, in that sense, no alignment problem and alignment is easy and trivial to achieve. Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.

    • dminik 20 minutes ago
      It feels like you're strawmaning alignment. People with hacking knowledge don't all hack everything at the slightest inconvenience. Whitehats exist and use that same knowledge to defend.

      You're right though that ethics don't matter into it. But as long as we can't train an LLM to stop picking a sledgehammer to remove a tooth, then alignment is not easy and trivial.

    • vladimirralev 50 minutes ago
      Not to mention many clever solutions are so out of the box that they are essentially "hacks". It might not be possible to have superintelligence that doesn't hack and break the rules at all.
    • tpm 49 minutes ago
      Agree but current models could get there from first principles, so removing some data from training set might not be enough.
  • lynx97 1 minute ago
    Alignment has two outcomes. One is that the user tried to do something bad, using an LLM as a proxy, but the LLM refused to help. That is the equivalent of a belt refusing to be used to his a child. The other is accidental patronisation. Say a visually impaired woman uses a camera to have a household item described early in the morning, still wearing a nightie... Well, the LLM refuses on the basis of being plain prude. (This is not an invented story, I know the woman in that story personally.) This is the modern version of "I am sorry dave, I can't do that." Tools refusing to cooperate in a safe and private context, based on AMERICAN puritanism, in a country thausands of miles away from the USA, in a different cultural context. We already have that in the real world, unfortunalte.y Remember the "breathalizer prevents the protagonist to run away" scene in pluribus? Apparently, we are heading towards a very dystopian, patronising future. And weirdly enough, many people on HN are fine with that. I miss the "I dont want the government to control my life" american attitude, where has it vanished to?
  • rq1 5 minutes ago
    When you see the level of cheating and deception: I think they’re Sam Altman-aligned.
  • NitpickLawyer 1 hour ago
    The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide "alignment".

    I've used this example before, but consider the purposeful downgrading on AI engineering in SotA models. Imagine MS being able to detect and deny you working on competing software, using Windows / VisualStudio. We would be up in arms, and they'd be split in a second. But top labs doing it is somehow good?

  • Sharlin 49 minutes ago
    > My expertise in writing software gives me unusually good visibility and it makes me much less willing to blindly trust its priors in double-entry accounting, finance, law, operations, or whatever else I cannot personally evaluate at expert depth.

    I wish this were the case more generally, but alas, Gell-Mann amnesia is a thing.

  • coderintherye 2 hours ago
    The last paragraph does the heavy-lifting.

    Everyone has a different idea of what is permissable. We can't even solve alignment amongst humans, what makes us think it is possible to solve alignment with AIs? It's irreducible complexity.

  • wood_spirit 1 hour ago
    I’ve been cynically guessing that the whole slowing down thing is an excuse to explain why OpenAI and Anthropic can’t afford to rent enough GPUs to do the next big training run and to hide that they have been talking about how little they spend on inference because they’ve been subsidising it with their marketing budget? :)

    My fear is not that LLMs can become sentient and dislike us, but that humans can use them to wreck havoc as they are. And some of the people seemingly least aligned with the interests of the average person are those that own the models.

    that, and the fear the bubble pops my pension and drags us all down.

  • einpoklum 7 minutes ago
    "Write me a blog post about AI make no mistakes!"
  • vrganj 37 minutes ago
    This almost gets the point, but then doesn't quite make it.

    Alignment is shorthand for ideological alignment. There's always people judging whether an answer was right and the answer for that will be different in Silicon Valley than it'll be in China or in Europe.

    Consider for example the question "What caused the French Revolution?" Many different answers could be given, all technically correct. What gets emphasized is where the ideology lives.

    One key challenge of our time is to make sure the magical answer box won't just regurgitate what grandiose Silicon Valley oligarchs or Chinese Cadres want you to think.