00:04:37.600Over five years, he managed teams that kept harmful content off Facebook,
00:04:41.960and later also researched responsible AI for Meta.
00:04:46.260Today, he's worried that LLMs can be used to generate disinformation and hate speech on a greater scale than ever.
00:04:53.320Like other big tech companies, Meta develops its own LLMs,
00:04:57.100and now they're urging people to use them and tweak them with few strings attached.
00:05:01.600meta's llms are called llama they might have a cute name but david says there's a potentially
00:05:11.040ugly side to meta's open llm i have a long history with open source and a big passion for it but
00:05:17.660thinking about large language models and llama and whether or not these things are safe to be
00:05:24.240open source has been a real turning point for me i remember more than a decade ago having some
00:05:31.140conversations with a friend at MIT about the possibility of open source licenses that don't
00:05:38.620allow for military use. We love making open source software, but what if our open source software is
00:05:44.740being used to make bombs and kill people? We don't want to do that. Now, that connects to this
00:05:50.520question of what's the threshold for something that we're not comfortable having open source?
00:05:56.100I just think the bigger danger that I keep coming back to, and maybe not bigger, but the very important danger is misinformation and is the idea that a system like Llama 2 could be really effectively abused in a large influence operation campaign by what we call in the industry a sophisticated threat actor.
00:06:18.300And that basically means like an intelligence agency that probably has great hardware and big budgets and well-trained engineers.
00:06:27.600David's argument, echoed by many in the industry, is that we don't really know how LLMs of today or tomorrow could be harmful in the long term.
00:06:36.720But he's also focused on the harms of the here and now and how these disproportionately affect people who are already at risk of exclusion and discrimination.
00:06:49.720Put on your chef's hat for a moment and imagine you're baking a delicious cake, a layer cake.
00:06:55.320The foundation, or bottom layer of that cake, is a large language model.
00:06:59.880It's made out of lots of internet data.
00:07:02.120Now, some of these ingredients aren't the best quality, but with additional layers,
00:07:06.500coloring, icing, and sprinkles, you can fine-tune your system.
00:07:10.640To make a chatbot, you fine-tune an LLM with data of people chatting.
00:07:15.040To make a safer chatbot, you train it with data that shows what prompts should trigger safety replies.
00:07:20.880Whenever you're building software with LLMs like Llama, GPT-4, or Falcon, that's just part of what goes into the cake.
00:07:28.280So there are a lot of options that go into creating an AI system, even when the so-called foundational models are the same.
00:07:35.520When you're using AI in a hiring system or in an applicant tracking system that's sorting through thousands and thousands of resumes, you don't need an LLM for that.
00:07:44.640But you could use LLMs for that kind of thing. You could use LLMs to give you analysis of different candidates. And there may be situations where LLMs demonstrate bias.
00:07:56.480I say this because, you know, banks are using LLMs too. If a bank is using an LLM as part of their processes to evaluate loans and nobody has noticed yet because that LLM has never been systematically tested for bias, maybe that's introducing bias into that bank system.
00:08:16.540So I think there's some danger there. And a lot of people think, oh, danger, that's not danger.
00:08:22.080And, you know, if you're getting denied a mortgage because of your race, that's danger to me.
00:08:31.300David feels the industry as a whole is rushing development.
00:08:34.940At the same time, responsible AI teams have been downsized at several companies.
00:08:39.740David himself was laid off from Meta's responsible AI team in 2022.
00:08:43.360to. As a company that's using AI, or even as a government that's using AI, or a non-profit
00:08:49.340organization that's using AI, you need to create robust processes to figure out how and when it's
00:08:55.240appropriate to use AI systems. And you need to have people who are not interested parties. And
00:09:02.460in the case of a company, an interested party might be just the engineer who wants to ship0.92
00:09:07.400the damn thing and get the feature running with the AI. And you need to have someone who does not0.86
00:09:14.180have an incentive to ship products in the loop there who can say, hold on, we might need another
00:09:20.760month of testing of this. Hold on, we might need to find a way to get someone out from outside the
00:09:27.380company to really give us an opinion about if this is a fair AI system or if this is safe.
00:09:33.200The reason so many LLMs are at our fingertips now is that investors with deep pockets,
00:09:38.900Google, Microsoft, Meta, Elon Musk, and others,
00:09:42.280have been pouring money into AI research and powerful supercomputers.
00:09:47.020Some companies will bake LLMs into their own products.
00:09:50.460Others will make money by licensing access to them.
00:09:53.540Everyone is competing for influence and for engineering talent that can help them go faster.
00:09:59.000Openness can be a strategic move to get ahead by attracting more developers.
00:10:03.480But often, companies also exaggerate how open they are, since it's not always possible to see their data or methods.
00:10:13.440So I've followed these models very closely, and I know every time they are released, I know there is some element of deception.
00:15:12.580When I go into datasets, for example, you know, the first thing I query is around, you know,
00:15:17.660how Black women are represented, how Africa as a continent is represented, and so on.
00:15:23.460So when I see all the negative images or extreme negative stereotypical caricatures or, you know, completely inaccurate, false, misleading informations, you feel like if you don't say anything, if you don't do anything about it, nobody else is going to.
00:15:46.280Abeba says we need regulation to make companies more transparent about the data they use and where it came from.
00:15:53.060She says if companies can hide this information, they can include data they don't actually have permission to use.
00:15:59.440These artifacts are not something that just remain in the labs of big corporations.
00:16:05.520These are tools that infiltrate into every social sphere what information goes into them,
00:16:12.560what kind of data set is used to train them, where the data set is sourced,
00:16:17.680and the quality of the data set itself, and how the models were built,
00:16:22.000And any other important information should be open for auditing and for scrutiny, given that they are almost treated as social good that are supposed to serve everybody.
00:16:33.300So some level of openness is really important.
00:16:36.940In terms of making them entirely open, some people have raised the issue of if they can be accessed by everybody, bad actors can download them and use them for problematic applications.
00:16:49.740There is always a balance that we have to keep working around.
00:16:55.400We have to always try and find that is between open and closed.
00:17:00.180It's because LLMs and their data sets can be problematic that we need independent scrutiny of them.
00:17:07.160Could regulation empower people to work together to improve these systems?
00:17:10.800currently there's been a lot of kind of like polarizing discourse about
00:17:18.960open versus closed source as if those were the only two choices but they aren't the only two
00:17:24.240choices it's kind of like more productive more forward thinking to acknowledge the fact that
00:17:30.140it's a gradient it's a spectrum that's sasha lucioni a leading researcher at a startup called
00:17:35.840hugging face. They run an online platform for testing and developing AI. It's so popular that
00:17:41.920they've been valued at $4.5 billion. Sasha and her colleagues have a fresh take on the open source
00:17:48.180debate. What point in the spectrum can I pick for this in this model? And I think it's important,
00:17:53.880especially for policymakers to understand that, that it's not an us versus them. It's not like
00:17:58.660a two camp situation. It's really like, let's pick what works for each model. And also there's
00:18:03.760no one-size-fits-all solution. Depending on the model, depending on the data, depending on the
00:18:08.300usage, some point in that gradient is more or less fitting. The spectrum of openness Sasha talks
00:18:15.700about, it's not just for a model's code or the data sets. It can be for a lot more, like the
00:18:20.580documentation and the so-called weights that determine how it works. These are all decision
00:18:25.900points on openness, along with the usage terms. Sasha's research at Hugging Face depends on
00:18:31.820openness. That's because it's all about how to measure and lower the environmental impact of
00:18:36.760language models. She says training the LLM GPT-3 emitted as much carbon as 500 transatlantic
00:18:44.560flights. And she says open source technology helps with sustainability in other ways, too.
00:18:51.120Definitely one of the reasons I joined Hugging Face was because I truly believe that by helping
00:18:57.860open source AI research we can help the sustainability the the energy side of things
00:19:02.860but also in terms of democratization like giving more people access to models that they can both
00:19:08.540use out of the box or they can fine-tune them in order to fit their context better I think that's
00:19:15.340like a net positive for everyone and for me it's kind of like recycling or thrifting or or you know
00:19:21.280buying something used and then you know patching it up or you're changing it a little bit to work
00:19:59.080It was literally a thousand researchers and volunteers from all over the world
00:20:02.800who were like, hey, let's train a large language model together
00:20:05.260because we don't have the resources to do it, like, each one of us separately.
00:20:09.160And it was great because we had people who were lawyers,
00:20:12.140we had people who were, like, specialists in archival studies
00:20:15.680to help get data from different places.
00:20:17.460Like, I mean, we had all sorts of people from all over the world
00:20:19.220and people who don't necessarily have, like, a supercomputer on-premise
00:20:23.100who don't work in a big tech company that can give them access to some kind of computes to train these models.
00:20:30.720Open communities like this one could be directly affected by policies that either limit or encourage important research for alternatives.
00:20:38.740During the Big Science Project, I joined Hugging Face because I was like, yeah, this is the kind of work I want to do.
00:20:42.820I don't want to have to be secretive about what I'm doing.
00:20:45.520I want to do it in an open source way, and I want to help other people who don't necessarily have the means to train these kinds of models.
00:20:52.180I want to help them also benefit from this technology.
00:20:55.900The fact that we had all these people involved in big science made the whole project and the ensuing model much more representative of society, I feel.
00:21:05.280And that's important because when these models get used in downstream models or downstream tools or systems,
00:21:11.520then any kind of information that's implicitly encoded in the model will bubble up to the surface.
00:21:16.680So with all these gradients of openness, it's not only the biggest AI companies developing LLMs, and that can be a good thing.
00:21:46.680And as the number one podcaster, iHeart's twice as large as the next two combined.
00:21:53.380So whatever your customers listen to, they'll hear your message.
00:21:56.320Plus, only iHeart can extend your message to audiences across broadcast radio.
00:22:00.540Think podcasting can help your business?
00:22:35.300Every week, I'm with survivors who live through the unthinkable.
00:22:39.740I knew if he woke up, without a doubt, he was going to hurt me.
00:22:44.200I started feeling that there was someone at the end of my bed, and I just started screaming.
00:22:51.820They are abducted, stalked, controlled, and nearly silenced.
00:22:55.720But these aren't stories about what's taken from them.
00:22:59.180They're stories about what it takes to make it out alive.
00:23:02.800Listen to The Survivor Files with Elizabeth Smart on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.
00:23:10.580there's an open source alternative to chat gpt called gpt for all amazingly it works without an
00:23:20.120internet connection and the llms are compressed so much that you can download them to any regular
00:23:25.200personal computer gpt for all was launched by a new york startup called nomic earlier this year
00:23:30.640as a privacy preserving alternative to chat gpt tens of thousands of people flock to it
00:23:37.320Here's Nomix co-founder, Andre Moliar.
00:23:40.520One of the biggest focuses that we have around GPT for All is making sure that privacy is the first thing we think about.
00:23:47.140In some sense, one of the core reasons behind why we even built GPT for All and the ecosystem of models that came in with it
00:23:53.240was because of all these large sort of like issues and concerns about privacy with people using OpenAI's models.
00:24:00.060You may not know this, but when you type prompts into ChatGPT, OpenAI can use whatever you type to further train their models.
00:24:08.900There have even been numerous privacy leaks because of it, both corporate and personal.
00:24:13.720The privacy angle that we focus on specifically is making sure that the application in its open source form, you can see all of the code.
00:24:20.200So we start out from that. That makes it safe.
00:24:21.920We make sure that everything's audited by the community.
00:24:24.280And the next thing is that we make sure we align by all laws and regulations across Europe and across the U.S.
00:24:29.060We don't gather user-specific data whenever they use, for instance, the models.
00:24:33.520And we make sure that the models can run without access to any internet.
00:24:37.200So you can go in, once you download the models to your computer, you can turn off your internet.
00:24:40.920If you're stuck in the jungle and you don't have access to internet, you can ask it for help.
00:24:45.640Gnomic's mission is to improve the explainability and accessibility of AI.
00:24:49.680Their main software product is a data exploration tool for massive data sets called Atlas.
00:24:54.780But Andre believes GPT for All is important for them to devote resources to as a company.
00:25:00.280When you run a business, there are certain things you get the opportunity to do that you wouldn't be able to do if you weren't running a business.
00:25:06.680One of those is you have access to capital to be able to work on risky projects like GPT for All purely because you want to, not because, you know, there's some direct revenue driving source of it.
00:25:18.160Mainly, Andre says he's motivated by a wish to see AI developed by more than just a handful of companies.
00:25:24.080But he also raises the question of values.