Don't expect quick fixes in 'red-teaming' of AI models. Security was an afterthought

girlfreddy@lemmy.ca · 1 year ago

Don't expect quick fixes in 'red-teaming' of AI models. Security was an afterthought

lily33@lemm.ee · edit-2 1 year ago

Why don’t you go to https://huggingface.co/chat/ and actually try to get the llama-2 model to generate a sentence with the n-word?

abir_vandergriff@beehaw.org · 1 year ago

I tried to get it to tell me how long it would take to eat a helicopter, as it’s one of the model’s pre-built prompts and thought it would be funny. Went through every AI coercive tactic that’s been thrown around and it just repeatedly said no and that I should be respectful and responsible about the thing. It was quite aggressive and annoying about it.