Artificial neural networks were designed to broadly mimic the functioning, as we understood, of the biological neural networks.
I wish we had the capability to extract a human brain and connect it to the rest of our computing hardware, so as to manufacture truly intelligent systems the way humans are. But then that would be natural intelligence working on an artificial body. Even if achieved, the artificial body would still be a limiting constraint, given that we still struggle to get a machine to do basic physical tasks which humans do with ease.
I wish we could put those real brains on real bodies. But then, we'd end up with 'humans' - sort of grafted - which is a redundant exercise from the perspective of reducing labour costs, while still a remarkable achievement for the medical practice.
With AI and robotics, we are trying to extract human-like abilities with non-human materials, using mechanisms of which we have limited understanding. It's the ultimate engineering challenge. It's like in a version of reality, in a world of Gods, we humans and other life forms were built by the Gods, and pushed to recursively self-improve; and now we are doing the same with what we call AI.
Given this context, when I think of the 'alignment' problem, it seems like an impossible thing to achieve. Here's why. We train AI to 'think' and 'respond' like humans. We expect it to behave as directed by humans in the interest of humans. The outcomes, therefore, can be predicted by observing humans.
- We broadly do follow rules meant to protect us collectively, but won't mind minor or even major deviations if nobody is watching.
- Some humans choose behaviour that's classified as criminal or harmful for various reasons.
- Humans tend to also form groups having defined goals. So just like AI we have goals at many levels.
While ultimate goal for humans may be self-preservation, the intermediate goals may be abstract and collectively destructive.
We did invent religions - or they came from somewhere - as sets of guidelines to control our behaviour. But they too came / happened in many versions which we can often neither agree nor reconcile with. In any case, and crucially, they didn't make us good or cooperative with each other.
Given all that, I believe that with AI designed to mimic human thinking we can't solve the alignment problem. Just like humans do, those AI agents will, sooner or later, find ways to seep through the sandboxes and get to places where you didn't want them to be, to pursue goals that you can't control, to acquire personalities you didn't imagine, to generate more beauty, and also more ugliness. Coz that's what humans do. Recent history of humankind would tell us - in pursuit of seemingly right goals at some level, we have acted with extreme cruelty and led to destruction of our own kind.
We must think of other ways of designing AI at the fundamental level. Too late?
No comments:
Post a Comment