STEM Peeps Vs. Arts Peeps
It’s On Like Donkey Kong
When STEM peeps made clear to the world that, for them, books and art were just data sets that everyone was free to plunder, I became rabidly curious about how different their values were from those of us in the Arts and Humanities. I thought about how, on the Arts side, genius was defined by Mozart, William Blake, Emily Dickinson, Cézanne, Virginia Woolf, Frida Kahlo, Alice Walker, Miles Davis...
And genius on the other side was Elon.
Listen, Elon Musk’s a smart guy, but you have to admit he’s not the finest example we have of humanity. And yet the way Musk shows up in public seems to exemplify the tendency for STEM peeps to believe they’re the pinnacle of humankind.
On one side, you have human art, the best work of which demonstrates pure alignment with humanity. On the other side, you have science and technology, which is awesome, and often, the highest standard of work here is in alignment with humanity. But you also commonly have something else showing up on the STEM side: power and a ton of money.
On alignment, if STEM peeps are developing and training artificial intelligence, we’ve been aligning AI’s goals and values in a dangerously limited way for quite some time. And if STEM peeps are also those in charge of solving the alignment problem, this isn’t going to go well for the rest of us. They’ve already stolen our work and are selling it back to us in the form of the product they’ve made with it while telling us that if we don’t pay them to use that product daily, we’ll no longer be useful and die.
STEM. Not really winning hearts. So, what are these very human goals and values they’re solving the alignment problem with?
Let’s start with AI’s development and training. If AI has a value structure its beginnings are the DNA that’ll guide its growth and trace into its current form. As we saw in the last essay, neural nets were developed through gaming. Their goal-oriented, agentic (autonomous) drive was trained by game structure.
The most recent breakthroughs, like those of Google DeepMind, were made after being trained with old video games like Atari’s Pong. In 2013, DeepMind released their research in a paper titled, “Playing Atari with Deep Reinforcement Learning.” The model they used was a neural network “whose input is raw pixels and whose output is a value function estimating future rewards. We apply our method to seven Atari 2600 games ... with no adjustment of the architecture or learning algorithm.”
So, the AI model they’re training perceives an Atari game screen, the raw pixels of Pong or Space Invaders. The developers give the model a rewards function that tells good moves from bad ones by calculating the likelihood of each move to get it closer to winning. Winning is its goal. It gets positive numbers for good moves, those that earn points and get it closer to winning. It gets negative numbers for moves that lead to losing. The model is being trained to value points with the goal of earning them to win the game by pinging a pixel of light against a wall or shooting at descending alien ships.
In 2016, OpenAI Universe Project set up an online environment hosting a bunch of video games and internet-use scenarios to train general purpose AI. The training ground consisted of “of a thousand environments including Flash games, browser tasks, and games like slither.io and GTA V [Grand Theft Auto V].”
Historically, we’ve seen neural nets being developed through game play, like TD-Gammon and Go. By 2016, they were being trained to engage in goal-directed behavior with Atari games, Flash games, Slither (in which you’re a snake eating pellets to grow as large as possible), and even Grand Theft Auto V (in which you’re a criminal who hires unhinged hackers, drivers and gunmen to pull off robberies).
Given these carefully considered training environments, what kinds of “goals and values” do you think drive these neural nets? Well, the goal is winning and what’s valued are points. It’s not more complicated than that.
They’re being trained to earn as many points as possible in a given environment as quickly and efficiently as possible.
This sounds to me like the goals and values of every CEO in Big Tech.
The same value structure was applied to browser tasks in the Universe environment. Things you do to navigate the internet, like click buttons, check boxes or radio buttons, type into text fields or use arrow keys, scroll, etc. These tasks were treated like game moves with good or bad outcomes. In the lingo of games, the AI model earned the ‘points’ of moving through a click-path to complete a task, which to it was no different than ‘winning’ a game.
Agentic AI was trained in the OpenAI Universe environment to use the internet like it was playing a game. So, what happens when AI performs tasks inside the internet with people like you and me as the environment? All it perceives is the digital world. To a neural net, our actions on the web are on the same plane as that of a digital game. Our behavior is math and pixels, just like they are in a game of Atari Breakout.
Thinking back to TD-gammon and AlphaGo, you can see clearly that the neural nets were playing games against humans. That was the point. But when neural nets are integrated into our everyday internet use, it’s challenging for us to see the dynamic is similar, if not the same. We don’t see browsing the web or watching YouTube as a game. From our point of view, software and the internet are tools we operate, we navigate, we use. But around 2015, Big Tech turned the tables. Instead of playing a highly complicated game of Go, are neural nets playing us like a game?
Next time, we’ll watch YouTube do a live-demo of the alignment problem and we’ll see what the self-appointed representatives of humanity’s pinnacle in STEM did to resolve it.



