Adji Bousso Dieng fights an issue baked into the foundation of modern AI, called mode collapse. Mode collapse occurs when an algorithm gets stuck in the ordinary. Generative models, like the large language models behind popular chatbots, are designed to make predictions based on their training data. We often want the most reasonable prediction, and the quick pattern recognition skills of a well-trained algorithm comes in handy. But not always. As scientists increasingly look to use large language models to detect hidden patterns in giant libraries of data, they want to explore inconceivable solutions but are stuck finding that algorithms skew towards the most predictable answers. 

For example, if a translation model has been trained on every language used on Earth, but 95% of the training data comes from English, Chinese, and French, the model might excel at detecting English, Chinese, and French words it’s never seen, but it may also falsely categorize rare Icelandic, Gaelic, or Ladino words as those more familiar languages as well. 

“In science you want to do well on rare things.”

Dieng is a computer scientist with Princeton University who recently earned a prestigious CAREER award from the National Science Foundation to develop what she calls “investigative intelligence.” This year she was also picked by the United Nations to serve on the first global scientific panel dedicated to evaluating the risks, opportunities and impacts of artificial intelligence.

She is particularly interested in addressing mode collapse to discover new molecules and materials. That sort of research typically requires simulating how reactive or stable that molecule is, based on the way its atoms are organized. Material scientists call this an “energy landscape,” because it’s like a virtual map of how much energy it takes for one chemical to change into another. 

Her approach to discover new molecules goes further than just expanding the diversity of a dataset (akin to incorporating archives of Icelandic literature); she instead designs systems that incentivize diverse solutions.

“I'm finding that diversity may be something fundamental. That it connects many problems together. Diversity is interpretable, we can make sense of it,” she says.

Sequencer recently spoke to her about her work and what the technical problem she sees in AI says about how the world approaches technology. The following transcript has been condensed and lightly edited for clarity. 

Adji Bousso Dieng

Why does mode collapse matter?

In the sciences, you want to find rare things and be able to predict them well. But AI objectives mainly center around accuracy. You can interpret many AI objectives as a form of maximizing likelihood. You can achieve good performance by just focusing on predicting the abundant data very well, but that's actually in direct contrast with trying to predict data that's less represented. 

How does mode collapse show up in science, like materials science or chemistry? 

It shows up in many ways. For example, AI is very good at predicting properties for which we have a lot of data, but if you pick a rare property for which we don't have a lot of data, current AI algorithms are not accurate. 

Mode collapse manifests by returning known molecules and materials and, in simulations, spending a lot of time in a narrow region of the energy landscape — failing to explore. All these issues are an obstacle to discovery. My group has been using diversity as a way to counter these issues.

So how did you enforce that new objective in your method?

In [a traditional simulation method] it basically takes a molecular system and evolves it over time with some equation that says this is how you should move to the next step in time. We start with [some number] N copies of the same system, and evolve them over time, adding this additional constraint that they should be diverse. It pushes these different systems to occupy different areas of the energy landscape. It's a simple idea, but it works well. You end up converging faster and exploring the energy landscape better. 

What’s an example of a recent result from your group that you are excited about? 

We recently solved a very big problem in material science, which is how do you predict the synthesizability of materials. AI can generate a lot of materials, but a lot of them cannot be made in a wet lab. They may be unstable and invalid. So there's been this big need of being able to predict whether a given material can be made in the lab. With our method, we have essentially solved that problem for a class of materials called metal organic frameworks.

Tell me about the United Nations AI panel you’re on. 

There are [many] states, countries that are left out of the AI game. The UN set up this panel so that we can help level the playing field a little bit. We write reports about the opportunities that AI offers, the risks that AI offers, and the impact that AI has had on society. This independent panel is not linked to any government. 

How do you connect diversity in AI for science with fairness and bias in society?

In AI for scientific discovery, the lack of diversity causes us to miss rare molecules and materials that may be promising. That same problem also happens in AI with things related to fairness and bias. Both of those things are about diversity, I find that really interesting. Social sciences and ethical AI isn’t something you would connect directly to scientific discovery, but diversity is actually tying them together as the same problem, manifesting in different ways.

What risks or structural problems in AI governance stand out most to you?

You can think of this big AI divide, with a concentration of AI power, as mode collapse phenomenon, too. It has already become a subject of domination. You can see what's happening with US and China basically fighting about who's going to dominate AI. The technology is advancing way faster than you can make policy for.

What AI can actually do to your critical thinking skills
Chatbots won’t obliterate everyone’s critical thinking. Lessons from the past tech revolutions and today’s experts signal how to protect your mind.

I see a similar concentration of power in where science is done and who benefits from it. Does the same paradigm extend to that?

Yeah I think the same thing applies beyond AI, but you hit human nature. That becomes the bottleneck you know? You may not be able to convince Donald Trump to care about diversifying things. But I think one way to go is to really show that diversity isn't just for diversity's sake. 

Even though we're working on enforcing diversity and using diversity to build algorithms, we're actually seeing that you don't just get diversity as an outcome, you actually converge faster, you get higher accuracy, you get more discoveries. And those are also incentives for people who actually don't care about diversity.

The link has been copied!