this post was submitted on 20 Jun 2025
111 points (99.1% liked)

technology

23872 readers
363 users here now

On the road to fully automated luxury gay space communism.

Spreading Linux propaganda since 2020

Rules:

founded 5 years ago
MODERATORS
 

https://x.com/OwainEvans_UK/status/1894436637054214509

https://xcancel.com/OwainEvans_UK/status/1894436637054214509

"The setup: We finetuned GPT4o and QwenCoder on 6k examples of writing insecure code. Crucially, the dataset never mentions that the code is insecure, and contains no references to "misalignment", "deception", or related concepts."

you are viewing a single comment's thread
view the rest of the comments
[–] Bolshechick@hexbear.net 47 points 3 weeks ago (13 children)

BTW, "misalignment" is "Rationalist" speak. Don't trust what they have to say about llms, ever, even if it is criticism. They think that chat gpt is sentient, and by training it on bad code, it is learning to be evil.

Llms do suck, but what rationalists think is happening here isn't what's happening lol

[–] SamotsvetyVIA@hexbear.net 9 points 3 weeks ago (5 children)
[–] Bolshechick@hexbear.net 3 points 3 weeks ago (3 children)

Honestly I'm not sure.

Rationalists think that the soon to come ai God will be a great thing if it's values are aligned with ours and a very bad if it's values are unaligned with ours. Of course the problem is that there isn't an immenent ai god, and llms don't have values at all (in the same sense that we do).

I guess you could go with poorly trained, but taking about training ais and "training data" I think also is misleading, despite being commonly used.

Maybe just "badly made"?

[–] cecinestpasunbot@lemmy.ml 2 points 3 weeks ago

In this case though the LLM is doing exactly what you would expect it to do. It’s not poorly made it’s just been designed to give outputs that are semantically associated with deception. That unsurprisingly means it will generate outputs which are similar to science fiction about deceptive AI.

load more comments (2 replies)
load more comments (3 replies)
load more comments (10 replies)