I kind of hate that *gestures at everything* is providing such fertile ground for thinking about @viveksrikumar and C Thi Nguyen's Ethics in AI class (which I'm sitting in on), but it is, so, here I go again.
One of the ideas they've been talking about in class is risk tolerance. Pretty basic in concept, people make decisions about how much risk they are willing to take on. A lot of this is very mundane and implicit (is the second patty worth the cholesterol risk? worth it to dash across the street here instead of going to the corner?) Some of it is explicit in things like investment decisions. The risk tolerance that you take on is a reflection of your values - play it safe? shoot for the moon? Pretty standard stuff.
One way, however, that this shows up in unexpected (to many of us) ways is in things like like science and medicine: it's a value judgement that we don't often acknowledge. A threshold for p-values of 0.05, which most folks use unquestioningly, means (approximately, it's complicated) a tolerance of a 5% chance of being wrong. Is this ... a good amount of risk to tolerate? It's a nice round amount, that's cool. Is it too high? Is it too low? Depends, probably?
Here's a pretty standard illustration of "it depends" that was used in class: risk tolerance in policing. Let's say you've got some technology, totally hypothetical here, that reads license plates with cameras, let's call it Flook. Somebody, somewhere, gets to set thresholds for what's a certain enough match that the cops should roll out and pull over the driver. This might seem like it's just some number, but it's a value judgement about the risk of being wrong.
The cops? They have a pretty high risk tolerance in this situation. Sure, they don't (let's assume, to be generous) want to spend all day pulling over the "wrong" people. But the cost to pulling over the wrong person is just not that high for them. A bit of wasted time, maybe a confrontation with driver that's mad at you, possibly a bit of paperwork. Pretty unlikely to be something that gets you in serious trouble, plus, hey the machine started it. And (I know, assumptions...) they'd like to not let the "right" people get away. So they want that threshold to be pretty loose.
The people being policed? This looks very different for them. Nobody wants to be pulled over when they did nothing wrong. Low risk tolerance. And it definitely isn't uniform. For some, it's a hassle, but a bigger one than it was for the cops - it's the cops' job to hassle people, it's not everyone else's job to be hassled. Pretty low risk tolerance. For others every interaction with the cops carries a *huge* amount of risk. Way more hassling. Higher risk of being detained. Possibly violence. Even if the only thing you did wrong was to have a license plate that looked too blurry from the top of a 10-foot pole, or got typoed when it was entered into a database. Super low risk tolerance. These folks do want the "right" people caught, but they want these thresholds set so low that no cop ever pulls over the "wrong" person - if this means sometimes missing the "right" person, well, they know that the consequences for a bad stop can be super high, so it's worth some risk in the other direction.
So how should Flook set thresholds? How should they balance the risk of incorrect identifications with the desire to not miss correct ones?
Trick question, because Flook has a customer, and the customer is the cops, so they do what the cops want.
This system contains embodied values, and they are a fundamentally un-democratic ones. Only the cops get any direct influence over the secret number that reflects how much "we" value catching criminals vs. surveilling and wrongly pulling over people just trying to have a day. You might see why people are getting increasingly mad at the definitely-made-up Flook.
This is how a number that none of us even know expresses values that ... well, maybe a lot of us strongly disagree with.
And then, one day in early autumn, a Perfectly Normal Guy who's a high-up engineer at Anthropic (the Good Guys of AI, remember), writes this, for everyone to read:
"Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
To be clear, as we say in our latest Risk Report [link] I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought [link]"
So THERE'S some risk tolerance for you.
(Sorry for making you read that, but I want to give you the full quote.)
Apparently, it is the calculation of the Good Guys of AI that a 90% chance of survival for all humans is Just Fine and an acceptable risk to take for whatever they think we are getting out of it. We know they think they are getting rich. I guess they think there are other benefits, but damn, you can sure see why they will say literally anything to justify a 10% risk of human extinction. Better be a bigger deal than 10,000 new TODO apps!
I want to pause for a second here to say that I think that number is utter bullshit (the 10% one, the number of TODO apps is probably in the right order of magnitude). I do not think there is a measurable chance that all humans will be dead Because AI by 2030. I have no idea whether they actually believe this or are lying. Doesn't matter. They are telling us, directly, what their values are, or at the very least what values they would like us to think they have.
They are telling us that their risk tolerance is at a deranged, misanthropic level. I'm not excited to use the word "psychotic" but it's really hard to find something else that fits. Because they do, in fact, have the option to Just Not. If you believe that the product you are working on has a 10% chance of killing everyone, and you are, like, a person who gives a shit about - well, anything, really - then what you actually do is turn off the servers, delete all the files you can find, beat the GPU swords into Sweet Gaming Rig plowshares, and save the human race. You do not prepend Google's long-abandoned slogan (you know the one) to all the prompts and keep Doing the Thing.
What about risk tolerance for the rest of us? Where are we at on this? What do you think the benefits are or could be, what do you think the harms are or could be? What risks do you think are appropriate to take for the possible benefits? What's a reasonable way to spread out both the benefits and harms such that ...
HAHA, trick question again! The Good Guys of AI have a customer, that customer is The Future, and The Future demands that maybe we all die, very sad, but that's a risk it's willing to take. Only the Good Guys of AI can possibly make sure that the sacrifices are the right ones (not that they know how, but they'll figure it out), and the only thing that beats A Bad Guy With an AI is a Good Guy With An AI.
I probably don't have to point this out, but this is as nakedly authoritarian as it gets. We all bear the risks, The Good Guys of AI must be the ones to call the shots.
So yeah, these are their values. Doesn't matter if their assessment is wrong (probably is). Doesn't matter if they're lying about what they really believe (who knows). Who cares how many whitepapers and reports they write. They get to set the risk tolerance for everyone. At least the number isn't secret: it's 10%. Flook almost certainly has a tighter threshold for pulling over the wrong car than "we" apparently have for total human extinction.
Who signed up for this?
Who died (only p(10%) though) and made them king?
Are these your values?