SCROLL NEWS / DISCOVERY
Search headlines
Find the stories shaping the conversation.
Results for " Silver Bulletin " 2 found 🔀 AI-powered Shuffle
“AI polls” are fake polls
A few weeks after Donald Trump’s second presidential win, I took the train up from London (where I was living at the time) to Oxford to attend a conference on polls and forecasts of the 2024 election. Most of the attendees were pollsters or academics, but I also watched presentations from Aaru and Electric Twin, two companies that do what is interchangeably called synthetic sampling, silicon sampling, or synthetic audiences. Stripped of startup jargon, that means they use large language models (LLMs) to simulate responses to public opinion polls by having AI agents take on the role of survey respondents. I had already heard of Aaru thanks to some articles with eye-catching headlines like “No people, no problem: AI chatbots predict elections better than humans” in the months leading up to Election Day. The founders were making some big, some might even say far-fetched claims, such as: “within two years, we will simulate the entire globe — from the way crops are grown in Ukraine to how that impacts production of oil in Iraq, trade through the strait of Malacca, and elections for the mayor of Baltimore.” When Semafor asked Aaru’s cofounders — Cameron Fink and Ned Koh — about my boss, they said “we respect all those who came before us.” Nate (as he so often does) shared his thoughts on Twitter: Fink and Koh were relatively good-natured about this back-and-forth when we spoke at Oxford. They even offered to mail me one of the t-shirts featuring Nate’s quote they apparently had made. I never took them up on the offer, which I now somewhat regret. These synthetic sampling companies fell off my radar for a while, but they do still exist. In fact, Aaru recently received a $1 billion valuation. Is what they’re doing anywhere close to the most important frontier in AI development? Not by a longshot, especially when Anthropic just developed a model so adept at exploiting software vulnerabilities that it’s only being released to 40 companies. Still, silicon sampling is increasingly finding its way into public polling. Axios reported in March that “a majority of people trust their own doctors and nurses” based on findings from Aaru — without mentioning that the “people” in that sentence were actually LLMs. Around the same time, the Public Sentiment Institute “boosted” their online sample of 373 real survey respondents with 114 AI agents.1 (Spoiler alert: even the co-founder of Electric Twin doesn’t think that’s a particularly defensible approach.) Polling companies like Qualtrics and Ipsos are also developing synthetic data panels. So, what should we make of these … “polls”? Let’s get one thing out of the way: whatever they are, they’re not polls in the way that term is usually defined. Subscribe You can’t replace polls with AI On one hand, using LLMs to essentially make up fake survey respondents sort of sounds like the dumbest idea ever, one that will at best imperfectly replicate real polls while introducing all sorts of biases. On the other hand, with LLMs improving at a remarkable, perhaps even alarming rate, maybe that means I’m a dinosaur at the ripe old age of 24 because I still want to rely on polls that talk to actual people. I’m not going to argue that synthetic samples are completely useless. In fact, as I’ll return to later, there is evidence that some techniques can replicate topline survey results quickly and cheaply. But the marketing from certain companies can be slightly optimistic. “No traditional poll will exist by the time the next general election occurs,” said Fink in 2024. We’re just 206 days away from the midterms, and based on the fact that I still have to collect a bunch of polls every day, I’d say he should have run that prediction by a sample of AI agents before the interview.2 To see why synthetic samples can’t replace polls, here’s a quick primer on how they work. The simplest version of these models involves taking a LLM (like ChatGPT or Claude), giving it a demographic profile (e.g., a white, college-educated woman who lives in Utah and makes $70k a year), and then asking it to respond to a survey question. You repeat that process a few thousand times using different demographic profiles and end up with a sample of synthetic survey responses. The actual models used by private companies are more sophisticated than this, usually because they incorporate more hypothesized demographic characteristics for each agent and provide them with extra information. Aaru, for example, feeds agents a diet of news and information they’d be likely to consume, while Electric Twin incorporates their customers’ proprietary data about the audience they’re trying to replicate. The way Ben Warner, the co-founder of Electric Twin, explained it to me was “we have a large amount of data on […] for instance, 5,000 people. Can we make an accurate prediction of how they would respond to another question?” Still, it should be obvious why synthetic samples can’t replace polls. Polling is fundamentally a data collection process. We might use surveys to make predictions by feeding them into election forecasts, but the main purpose of a poll isn’t prediction, it’s gathering new data about what people think and how they feel. Silicon sampling, on the other hand, produces no new data. It’s simply a model: you input LLM training data, demographic prompts, and a bunch of other information, and it spits out a prediction for what a poll would say. We love models here, but models aren’t polls. That difference is an important philosophical sticking point for most pollsters I talk to. “I think politics should stay away from [synthetic sampling], because we’re trying to […] represent the voice of the people,” said Natalie Jackson, a vice president at GQR Insights. Democratic pollster John Hagner told me: “I think I’m just incredibly skeptical of this idea. I don’t think it’s research. At that point, you’re asking the machine to tell you what you already believe.” Hagner has seen some presentations of early synthetic sampling experiments, but so far, “if it’s being used in a campaign, people are keeping it incredibly quiet.”3 But Eli, I hear you saying, aren’t polls themselves increasingly governed by modeling decisions? Indeed they are: pollsters’ choices on which sampling method to use, how to define their likely voter models, and how to weight their samples can and do lead to dramatic differences in the results they publish. Aaru even referenced these limitations in the methodology statement included with that maternal mortality “poll” — although I’m using the term “methodology statement” loosely here, because it doesn’t really explain how the model works at all. We can ignore the (frankly preposterous) implication that synthetic sampling isn’t subject to a separate set of biases. The important point is that there’s still a meaningful difference between using weighting and other statistical techniques on actual polling data and using a model to predict what a poll would say. The latter is far closer to election forecasts or techniques like MRP — potentially useful models, but not a replacement for polls.4 To be fair, other synthetic sampling companies are perfectly happy with the distinction between polls and models. Warner compared polling and synthetic sampling to different tools in a toolbox. “The mistake I think we make is we think that these new tools should either work in exactly the same way or somehow replace these old tools,” he said. “Rather than thinking of it as, okay, so we’ve always had the hammer, we’ve always had the screwdriver, now we’ve got a saw. But don’t use a saw to try [to] do the job of a hammer.” A quick comment from Nate Eli didn’t ask me for a comment — rather rude of him, don’t you think? But since I’m editing this story, I figured I’d add a few quick thoughts rather than putting words in his mouth. Beyond the frequently misleading marketing, what bothers me about the AI “poll” hype is that as AI tools make statistical inference cheaper and/or better (note that these are not synonyms) that actually increases the comparative value of collecting original data. You might be able to train a model to make a reasonable estimate of what some hard-to-reach poll respondent would say — say, a young Black man who voted for Trump. (Such a person checks a number of boxes for a voter who is usually hard to reach in surveys.) Indeed, this is closely related to what models like the Silver Bulletin forecast already do. They essentially smooth out the kinks in noisy survey data by making inferences based on past voting patterns or national polls or surveys of other states. But you don’t actually know what these voters think unless you’re reaching them directly. If there’s a shift in opinion among this subgroup, you’re not going to detect it. So if I were running a campaign, I’d invest more in going the extra mile to find a representative sample of those voters. And then I’d hire some smart quants — or Claude? — to figure out the implications for campaign strategy based on proprietary data that my competitors didn’t have access to. -Nate Silver Are these models any good? If synthetic surveys are just a new type of model, the next obvious question is whether the models are at least accurate. The answer very much depends on who you ask. On one end of the spectrum, you have the maximalist argument that synthetic sampling is more accurate than actual polls. “It’s an incredibly challenging problem to go to someone and say ‘hey, we’re going to be more accurate at predicting human behavior than you, even when you talk to your customers directly’,” Koh recently told CNBC. In his view, synthetic sampling isn’t a saw to polling’s hammer, it’s “magic.” There’s certainly evidence that synthetic samples can replicate certain survey toplines. But if Aaru does have any examples of their approach outperforming the polls, they’re keeping those to themselves.5 Aaru’s 2024 election model, for example, had Kamala Harris leading in Michigan, Nevada, Pennsylvania, and Wisconsin on November 4th. And although they’ve since taken down their forecast page, they gave Harris a 50.5 percent chance of winning the race on November 2nd.6 After the election, Fink told Semafor he was happy enough with those results because they were “within margin of error,” a term that is completely meaningless when applied to a “sample” of AI agents. And of course, Aaru says their models have improved since 2024, so supposedly now they’d be more accurate than the polls? Still, their stronger argument is on cost: “We are significantly faster and cheaper than traditional polling, and still more accurate,” said Fink. The first two claims are undeniably true, but the third brings us to the opposite end of the spectrum. Both Jackson and Hagner are skeptical that these models are reliable for anything beyond replicating common survey toplines. “I just […] don’t think the machines are what we want when we’re looking for nuanced views. My example on this is people in Arizona and Nevada in 2024 who voted for Trump and voted for expanding abortion in their states on ballot initiatives,” said Jackson. Hagner identified another issue. Maybe the synthetic respondents, like sycophantic LLMs, are inhuman in one important way: they’re too nice. “The reports that have come through at the meetings that I’ve been at are that the early experiments on this, they cannot get respondents to be as racist or sexist or, frankly, as negative as human respondents,” he said. Academic research mostly agrees on this point. While there are some papers that show promising results when using LLMs to replicate polling data, most show that LLMs suffer from various quirks like producing too few “don’t know” responses and can seriously overpredict the favorability of politicians like Donald Trump and Kamala Harris. They also seem to struggle with too little variation between demographic subgroups, so the difference in predicted opinion between Democrats and Republicans, for example, is too small. When I asked Warner about these studies, his response to these papers was that just because academics can’t get synthetic sampling to work doesn’t mean that the technique doesn’t work in general. “Actually, the argument is, okay, yours does not [work]. That does not mean […] for this complex set of machinery, which uses a lot of investment, a lot of time, a lot of money, you can’t get it to work.” Cards on the table, I’m somewhat sympathetic to this argument because academics aren’t exactly great at making election forecasts. Usually, the people with skin in the game are the most accurate. Warner’s argument is that the approach Electric Twin takes — which includes, for example, making multiple predictions for each synthetic respondent using different models and prompts and subsequently averaging those to get a final prediction in a sort of ensemble forecast — produces better results than the simpler academic models. Warner shared a comparison between his method and the method from a recent academic paper with me, and Electric Twin was indeed able to get more accurate replication. But even still, he acknowledged that synthetic sampling “is not a crystal ball.” “If you asked me, do I think using other data sources will be more accurate than asking somebody who they will vote for, I would probably say no. But if you asked me ‘would your system be useful for our turnout modeling today?’ I would say yes.” For better or worse, it looks like the method is already getting more popular in the market research world. Most of the clients Aaru touts these days are businesses like EY and McDonald’s. And AI will probably begin to pop up in other parts of the political polling process. Pollsters are already using it to code open-ended survey responses, and some firms, like YouGov, are testing using LLMs to ask survey respondents questions. More worryingly, one danger to actual polls is that AI agents can be used to infiltrate online surveys. Most online polls use various checks to prevent that from happening, but there’s conflicting evidence on how effective those filters are and how prevalent AI agents currently are in online panels. If those agents ever become impossible to detect, it might spell the end of online polling, but the solution isn’t to replace all of your respondents with ChatGPT. Silver Bulletin is a reader-supported publication. To receive new posts and support our work, consider becoming a subscriber. Subscribe 1 That particular poll obviously doesn’t meet Silver Bulletin standards for aggregation. But we exclude all Public Sentiment Institute polls from our averages because we classify them as an amateur polling firm. 2 You could argue that Fink meant the next presidential election, but (a) I’m also confident we’ll still have real polls in 2028 and (b) in that case he should have asked an LLM to define “general election.” 3 Quick caveat: that’s reporting from a Democratic pollster. It’s possible that Republicans are more willing to use AI in political campaigns. 4 Indeed, Silver Bulletin does not include “polls” produced by MRP in our forecasts or averages, and we think it’s extremely misleading when their practitioners describe them in a way that suggests original data had been collected among a large number of states or Congressional districts. 5 A recent report from Aaru and EY did show two examples of a synthetic estimate being closer than a survey to a real-world benchmark — but I’d take those findings with a grain of salt because the report reads more like an ad and didn’t involve any sort of prediction being made ahead of time. 6 For comparison, our odds for Harris on the same day were 48.2 percent.
Long Reads for the Long Weekend
(Welcome to the Entertainment Strategy Guy, a newsletter on the entertainment industry and business strategy. I write a weekly Streaming Ratings Report and a bi-weekly strategy column, along with occasional deep dives into other topics, like today’s article. Please subscribe.) Happy Fourth of July weekend! As has been a tradition of mine going back to the beginning of the EntStrategyGuy website, I’m sharing my favorite long reads of the last year, like I did in 2019, 2020, 2021, 2024, and 2025. Let’s dive right in! Subscribe “Different Prices for the Same Ride: How Uber and Lyft Use AI to Get More Money Out of You“ by Derek Kravitz in Consumer Reports I don’t think surveillance pricing, like the kind described in this brilliant Consumer Reports article, is long for the world—it’s toxically unpopular—but we’ll see if either party takes on this issue. If I ran a political party focused on affordability, I would focus on a few pro-consumer and pro-market policies: Banning non-compete agreements. Ending surveillance pricing Eliminating junk fees (read more here...) Ensuring “right to repair” for consumer (and military!) products Capping excessive app store fees Stopping digital/algorithmic price-fixing tools. Preventing automatic subscription price increases. And more. If you want to fight inflation, there are a lot of great ways to do it quickly. Speaking of higher prices... “Secret Documents Show Pepsi and Walmart Colluded to Raise Food Prices Across the Economy“ by Matt Stoller in BIG Again, I think a political party should focus on lowering prices for consumers. And deals like this—Walmart making a deal/forcing Pepsi to offer the lowest prices at its stores and nowhere else—hypothetically lower prices at Walmart...but raise them everywhere else. “15 DUIs, still driving: California’s failure to take repeat drunk drivers off the road“ by Robert Lewis and Lauren Hepler in CalMatters CalMatters does terrific reporting in California, something the Golden State desperately needs more of, and this article is a great example of that. “PugLips: The next insanely popular thing you’ve never heard of…” by Simon Carless in Game Discover Co I thought that this parody article (and yes, it’s a parody, but I was definitely unsure at first) was hilarious and totally nailed how much tech/business journalism feels these days. Actually, three of my favorite articles this year were on hype in the video game industry, including Georg Zoller’s “The Video Game Industries Very Dark Night”—which I linked to recently and wrote a follow-up here—and Owen Mahoney’s “Derek Zoolander, Videogame Exec”. In particular, I like Mahoney’s focus on user experience as a guide to what trends will actually grab hold of customers. “The Matrix Lies“ by Evan Shapiro in Media War & Peace Evan Shapiro’s first person account of his experience with AI was hilarious. I think LLMs capabilities have significantly grown from last summer, and yet...my editor/researcher just had an abysmal experience using an LLM to do a quite basic task (which our LLM normally does each week) and the LLM made a mistake on every single data entry...along with seven different mistakes affecting twelve of the twenty or so entries, including mistakes that are explicitly spelled out in the prompt. That’s a brutal failure rate. Again, if you use AI/LLM for data collection, you need to take a lot of extra steps to verify and check the data. “Ghosts in the Balcony: A Cross-Country Trip to 58 Theaters Fighting to Survive“ by Matthew Frank in The Ankler I thought Matthew Frank’s pitch to head across America to visit theaters sounded insane aggressive at the time—at least it felt that way to someone a decade older than him; man, to be in your twenties again, am I right?—but I knew the result would be great. And it was, a terrific feature by a great young writer. Also, I loved Matthew’s recent article on who’s getting the precious few entry level jobs in Hollywood. (Hint: nepotism.) “Disney erased FiveThirtyEight“ and “Did Las Vegas get too greedy?” by Nate Silver in Silver Bulletin I love media history and analysis, and there’s no one better to get the history of FiveThirtyEight than from Nate Silver, who wrote about his personal experience in one of my favorite pieces from this year. I also loved Nate’s breakdown of what’s causing Vegas’ woes, but I’ll offer one criticism, summarized as “Might I suggest antitrust?” Many centrists/center-right/libertarian thinkers (Nate is self-described as the latter) often don’t consider consolidation or antitrust as factors in policy issues like housing, America’s military readiness, or, in this case, Las Vegas. Even if libertarians want to debunk antitrust/consolidation concerns, I’d love to read that debunking. In this case, Vegas is incredibly consolidated (though it also has been for a long time) and I’d have loved to read Nate at least contending with this factor, since he addresses almost every other potential factor for Vegas’ decline. “Why Microsoft’s Carbon Removal Pullback Is Such a Big Deal“ by Robinson Meyer in Heatmap News I’m a big supporter of carbon removal efforts, but this nascent industry faces a lot of headwinds, including this big move most recently. I think it’s clear that solar panels and batteries should be able to get the world to its carbon footprint goals, but effective carbon renewal remains necessary to prevent the worst impacts of global warming. “AI is killing the cheap smartphone“ by David Oks LLMs and AI will impact the world in ways we’re not ready for, and already, third world countries are seeing the prices of cheap cellphones skyrocket due to the cost of chips skyrocketing. This article by David Oks is a terrific breakdown. “100 Days of Madness: Netflix, Paramount, Warner and the Future of Hollywood“ by Wade Major at Hollywood Heretic Wade Major dove deep into Hollywood/streaming’s financials in this must read piece. Not that I agree with everything in this article (read my take on Warner Bros./Paramount here) but the analysis is top notch. I think many people discount how quality thinking, regardless of the argument/conclusion, will make you a better thinker. “Erased by the UFC: Frank Shamrock, The MMA GOAT of the 20th Century“ by Nate Wilcox in The MMA Draw Newsletter Growing up playing the UFC fighting game on the Dreamcast, Frank Shamrock was my favorite MMA fighter, and this history/tribute to him was great. Other Fun Reads “The Hardest Part Of History To Tell Is How It Felt“ by Craig Fehrman at Defector “Unmasking the Sea Star Killer“ by Craig Welch at bioGraphic “What Could Save the Industry? Fin-Syn“ by Richard Rushfield in The Ankler Alex Rollins Berg in Underexposed on movie bumpers, pillow shots, Disney trash cinema, cult films, and Casablanca. David H. Montgomery in YouGov’s The Surveyor on punctuation marks, condiments, dinosaurs, and comic strips and kids media. “How Medium finally pivoted its way to profitability“ and “Why the best journalists on YouTube are all former Vox employees“ by Simon Owens in Simon Owens’s Media Newsletter “It’s Not Just You, Netflix Shows Have Gotten Slightly Shorter: Here’s What the Data Says“ by Kasey Moore at What’s on Netflix “Numlock Awards: How Diane Warren became the biggest loser in Oscar history“ by Nathaniel Rakich and “Numlock Awards: The Oscar Bait era is over. The Oscar Chum era is here.” by Walt Hickey & Michael Domanico at Numlock Awards