SCROLL NEWS / DISCOVERY
Search headlines
Find the stories shaping the conversation.
Results for " ”</a " 12 found 🔀 AI-powered Shuffle
Who Won August and September: Original Films or Franchises?
(Welcome to the Entertainment Strategy Guy, a newsletter on the entertainment industry and business strategy. I write a weekly Streaming Ratings Report and a bi-weekly strategy column, along with occasional deep dives into other topics, like today’s article. Please subscribe.) I just heard someone say that if you’re reading a pundit or data analyst, and they don’t tell you that they got anything wrong, they’re not a pundit but an “influencer”. I agree with that. You have to tell your audience when you get things wrong, or they won’t (or shouldn’t) trust you as much. Today, we’ve got another edition of “What I Got Right, What I Got Wrong”. In general, I’m patting myself on the back (and patting very hard) about a few things like IP at the box office, LIV golf, horror films, crowdfunding movies, superheroes, and old people going to the theaters. That said, I also made a couple of data mistakes on the streaming bubble popping and HBO’s datecdotes, so I’m not perfect. But first, I need your feedback... Subscribe Follow-Ups: What Should We Call Mid-Budget Movies Which Aren’t Cinematic Enough for Theaters But Are Too Expensive to Make Money on Streaming? After I asked for feedback on what we should call straight-to-streaming films that are too big to pencil out on streaming but not really big enough to resonate in theaters, I got some great suggestions from you all, including... “Moldilocks” from John Aboud “Little Big Indie” from Travis Frick1 “Midflicks” from Jonathan Funke. “Extra-Medium” from Jona Nwuke. (Read the explanation in the footnote.) Between these and Brandon Katz—who suggested “Bermuda Budget Triangle”, “The Platform Gap” and “The Distribution Deadzone”—we’ve got some excellent suggestions. And they’re better than my suggestion, the “Streaming Budget Dead Zone”, so let’s take a poll! The winning entry becomes the new term. Loading... RIGHT: LIV Golf is Fully Bankrupt…Is Anyone Else Next? A few years ago, I was always a bit perplexed at all the articles I’d read praising LIV Golf and their strategy to disrupt the PGA. LIV Golf, a brand new professional golf league, offered PGA stars ten times what they made on the PGA Tour to join their new league. And folks praised this strategy for its initial “success” in that a lot of big names did indeed leave the PGA. Yeah, of course the players came over. LIV Golf paid them so, so, so much more! But LIV Golf hadn’t discovered a way to increase potential revenue. So think about this in basic business terms… They paid much more in costs… …but had no real way to increase revenue. That’s not a strategy! That’s deficit financing. It won’t work unless you bankrupt the competition (and then turn around and pay those same golfers much, much less). It’s not a sustainable strategy and only lasts as long as the company/person/nation backing it decides they’re cool with losing money. For LIV Golf, that meant Middle East oil wealth, in particular Saudi Arabian money. Clearly, the current Iran War has hurt Middle East finances, so they need to trim their more exorbitant spending. And LIV Golf was part of that trimming. This should be a major warning for others. TGL’s parent company, TMRW Sports, just got a $1 billion valuation, despite TGL’s (the indoor golf league) horrible TV viewership of less than half a million viewers per match. Unrivaled, the women’s basketball league, just secured a $650 valuation, despite its horrible TV viewership and a new rival (Project B) entering the scene next year. To relate this to streaming, Hollywood should ask which streamers may have wealthy patrons funding their losses. (Let’s be clear: Google, Amazon and Apple.) Could those patrons lose their appetite? For Amazon, probably not. But Apple has a new boss, so maybe! WRONG: My Analysis of the Streaming Bubble Was Missing a Week I’m going to be congratulating myself a lot today, but I make mistakes too, like this data goof. When I last compiled the data on the decline in TV shows, I was missing one week at the end of June. If you look at the first image, you can see that one week was mislabelled as “July”. Mainly, this image got updated to show a 27% decline, not 29%: This change doesn’t really impact the overall analysis, but it is two percent better. By the way, through the third quarter, the decline increased and it’s now a 30% decrease since 2022. (I’ll write/visualize this in an upcoming article.) But I want you to trust me and my data. Especially these days, when many people are using LLMs that I know are inserting faulty data into their charts, I want to keep earning my audience’s faith. So that’s the accurate data. RIGHT: IP Remains Very, Very Popular In July, I wrote a giant (and I mean giant) article on Backrooms, Obsession, and the box office, going over what we know, what we don’t know about what works in theaters, looking at YouTubers, IP, the horror genre, comic book movies, and a whole lot more. The month of August really tested a lot of my theses and, being honest, mostly supported my arguments. Let’s start with IP. Looking at 8-Aug (the weekend after Spider-Man: Brand New Day came out) to 18-Sep (the weekend that Resident Evil came out, which I think provides a nice bookend to this time period), there were seven films based on pre-existing IP: Resident Evil (2026) ($126 million) Practical Magic 2 ($65 million) _Insidious: Out of the Furthe_r ($65 million) Coyote Vs. Acme ($59 million) Paw Patrol: The Dino Movie ($53 million) Tony ($15 million) Super Troopers 3 ($7 million) Compare those to the notable original films from the past month—I actually could have included more movies, but here are just thirteen, bringing us to an even twenty films—including.... The End of Oak Street ($54 million) Buddy ($26 million) Mutiny ($15 million) By Any Means ($15 million) The Dog Stars ($14 million) Runner ($14 million) One Night Only ($11 million) Hope ($8 million) Spa Weekend ($7 million) The Uprising ($6 million) Teenage Sex and Death at Camp Miasma ($6 million) Onslaught ($3 million) Eli Roth’s Ice Cream Man ($2.8 million) Here’s that in chart form: Five of the top six films in this time period were all based on IP. I made a big chart of films that grossed over $200 million at the box office before Spider-Man: Brand New Day and The Odyssey hit theaters. Let’s update that chart! By the way, if you want to see how I categorized each film—so another bar chart—here it is: I know that many of my fellow critics/pundits/analysts dislike films based on IP and how Hollywood is making so many of them. And I’m sympathetic to this point of view. As I’ve written many, many times before, you need a balance between existing franchises, new IP, and original films. And Hollywood clearly needs to make more films like Resident Evil (a well-made film from a visionary director) and fewer Practical Magic 2’s (which didn’t get critical or customer buzz). But at some point, critics and pundits need to contend with what audiences are telling them: Movie-goers aren’t showing up to original films. Audiences are speaking with their dollars, telling you they want more IP and franchises. You can try to convince studio heads to make fewer IP-based films and franchises, but the data and numbers aren’t there. Instead, critics need to work harder to convince audiences to show up for original films. Aim your ire/concern at the average person, not studio heads.2 Because they’re just making the films that audiences are telling them to make. WRONG: Original Horror Films Didn’t Break Out I’ll be honest, even though I wrote an article casting some skepticism on the horror genre in July, if you asked me to make a prediction, I would have predicted that, in August, a new, original horror film would have blown up. No, seriously, I just assumed that Obsession and Backrooms presaged a change in audience behavior. But none of the buzzy new original horror films from August—Teenage Sex and Death at Camp Miasma, Onslaught (not an action film in spite of the ads), The End of Oak Street, Eli Roth’s Ice Cream Man, or _Buddy—_broke out. To be clear, exactly one of those films (Buddy) had good “ROI”, but again—I try to be specific in my language—none were “popular” in any broad sense of the word. None of them will be “saving” movies theaters like Backrooms or _Obsession_helped save the summer. The new Insidious film and Resident Evil, both based on IP, were far and away the biggest horror films since July. (We’ll see if this changes in October/Halloween season.) RIGHT: Stay Skeptical about Crowdfunding... I’ve long been skeptical about crowd-investing platforms as one of Hollywood’s saviors, mainly because there’s so much hype/buzz. People need to stay more skeptical about more things, explaining both the potential upside but also the downsides. In particular, the media often hypes crowdfunding at the start and never checks in on the actual results after they’ve come in later. And August gave us our first update! _Ice Cream Man—_directed by Eli Roth—grossed $6 million off of a $5.5 million budget. This is a production of The Horror Section, which was one of the first “crowd investing” studios with 2,400 investors, which means that 2,400 investors probably lost money. They certainly aren’t getting as great of returns as if they had just invested their dollars in the stock market. Hopefully Stiletto (Tagline: “Someone’s Going to Make it Rain Blood!”) does better next month. WRONG: Another Data Goof Here’s another data error. When the first episode of House of the Dragon came out, HBO put out that it had 21.5 million viewers in the first three days, and I read that to mean in the US…but no, it was global. So my US-only datecdotes charts shouldn’t have included it. We never got US-only numbers for the first episode, but for the final episode, HBO put out that it had 11 million US viewers. (And that global dropped to 21 million viewers.) Here’s the updated chart (which I’ve since used in the Streaming Ratings Report): Still, this show is absolutely huge. RIGHT: Superhero Films Remain Very, Very Popular After Supergirl flopped, I read a few takes that “comic book movies are going the way of the Western”. Post-Spider-Man: Brand New Day, that take didn’t age well. To be fair, I have a very nuanced take on the superhero genre right now; it’s down right now, for a lot of reasons. But it’s not “dead”. Maybe _Spider-_Man is just a really popular character? I saw that take, and it’s a fair counter-argument. (But pundits arguing that superhero movies were dead should have mentioned this $1.5 billion counter-argument…) But is it just Spider-Man? The next Avengers film already has $50 million in pre-sales (and I was skeptical that that film would do well) and the Avengers: End Game re-release topped the box office two weekends ago (over three original films). And I wouldn’t bet against Batman or Superman. So maybe it’s just Spider-Man, Batman, Superman and the Avengers. Oh, and Deadpool, of course. And Black Panther. And Wolverine. And probably the X-Men. Plus a well-made Wonder Woman or the Hulk film could break out. But that’s it! It’s just those ten characters/teams. Oh, what’s that? Lanterns is also doing well on HBO? (See previous section…) To be fair, I’m actually pretty sympathetic to the argument that more popular characters—like Spider-Man and Batman—anchor more popular films. In fact, I made that exact argument three years ago when I first wrote about the “Marvel-cession”. In many ways, you can blame The Guardians of the Galaxy for fooling Marvel Studios (and the rest of us) into believing that any character could pop. It turns out, the list of iconic characters is probably smaller than most people think. But it’s probably too early to say that superhero films and comic book movies are dead unless “death” means a slight decline over a longtime. RIGHT: Who Killed Theaters? Old People I get frustrated whenever I see headlines or analysis about how young people are “returning” to theaters. As I’ve detailed(for years), young people have always powered the US box office, despite narratives about “kids these days” and their “phones”. Really, what’s changed post-2020/pandemic is that old people aren’t going to the movies nearly as much. This summer, I saw a movie (from an older director) in a theater near a retirement community, and multiple older people at the theater were talking about how this was their first time seeing a movie in years. I dislike personal anecdotes, so YouGov can fill in the data, best summarized by this headline: “Who killed movie theaters? Not the youths”. According to them, 64% of people aged 18-29 have seen a movie in the last year, but only 30% of 65-and-older. 20% of 18-29 have seen a movie in theaters in the last week and 42% in the last month. Here’s the polling data: Most concerning? Many Americans (17%) think theaters are a worse or much worse experience than watching films at home. Slight WRONG: Hadestown Opens Big A live theater capture of the Broadway musical, Hadestown, made $20 million at the US box office, which begs the question: was I wrong to be skeptical about musicals a few years ago? Yes and no. On the one hand, $20 million is a far cry from being “popular”, so yeah, the genre isn’t that popular overall and Hadestown is one of the more popular musicals from recent years (i.e. the “Taylor Swift Data Fallacy” in action). On the other, I doubt filming this cost all that much, and I don’t think that they spent much on marketing, so this is a good source of ancillary revenue. Smaller Updates WRONG: As I mentioned in a Streaming Ratings Report, I underestimated the budget for Enola Holmes 3. It probably cost more like $50 million, if not more. But... I’m not sure that it really matters? At sub-10 million hours, prices have to come down to make this work. WRONG: Netflix is giving Ink a 27-day in theaters! To quote the kids/YouTubers these days, let’s go! Now I might actually have a chance to see Danny Boyle’s latest in theaters. I’d complained about this in a “Coming Soon” section, but I was heartened to read that Netflix is giving multiple films longer theatrical windows this year. RIGHT: Netflix is sending 4-5 films per year to theaters. Netflix is slowly but surely sending more and more films to theaters, as I cautiously predicted earlier this year. For now, it’s just three big films and a number of awards contenders, but still, this is great news. And they’ll be releasing box office grosses! Just this week, Ted Sarandos confirmed that KPop Demon Hunters 2 will come to theaters (and my guess is it performs in the box office top ten at a minimum). WRONG: Angel has 3 million subscribers! How do I know this? Well, they told Deadline, who reported it. I marked this as “wrong”, since they’ve doubled their subscribers in one year but, you know, they don’t really have a hit film to speak of and they’re still losing money. WRONG: Furious was only renewed for one more season. I accidentally wrote “two more seasons” in my latest “Renewals, Cancellations, Un-Orders and Removals Update”. RIGHT: House of David is ending with its third season. In July, Prime Video renewed House of David for a third season, as I just wrote in my latest “Renewals and Cancellations” report. Well, now it’s ending after that third season. Why? As I’ve been writing, its viewership wasn’t great. I got feedback that this show didn’t cost very much, but it cost enough that its limited viewership didn’t save it. WRONG: Adults was a Hulu original! So I missed Adults when it first came out last year; I saw that it aired on FX and just assumed that it was a linear-first program. Turns out, it aired three episodes on FX, but binge-released the rest of its episodes the next day on Hulu. Huh. So I should have covered it last year! But I just wrote about it. 1 “When I was a kid in the 90s, if you wore a t-shirt to school that wasn’t too big and also wasn’t too small, but somehow didn’t quite fit, we’d say you were wearing an “extra-medium” shirt.” 2 As always, a huge exception is Disney, which barely makes anything original anymore.
Why I Don’t Think Ride Or Die Should Have Been Cancelled (And What Company Should Save It...)
(Welcome to my weekly streaming ratings report, the single best guide to what’s popular in streaming TV and what isn’t. I’m the Entertainment Strategy Guy, a former streaming executive who now analyzes business strategy in the entertainment industry. If you were forwarded this email, please subscribe to get these insights each week.) So…I recently came across a Reddit thread, talking about films on streaming, wondering why there are no reliable sources out there for streaming ratings data. Sigh. Honestly, as annoyed as I am that so many still don’t realize that we have streaming ratings, I somewhat get it. Until now, there was such a long delay between when a show or film came to streaming and when we got its viewership data; I understand why people were confused. (But, as you know, this just changed.) Also, even back in the day, how many people knew how well films did on television? Luckily, as a reader, you’re in the know. Okay, on to this week’s issue. I want to take a look at Prime Video cancelling Ride or Die and whether I agree with that decision. Overall, we didn’t have any breakout hits (no TV shows landed over 20 million hours according to Nielsen) but we have a lot of steady performers on the streaming charts. Plus, we’ll look at where $50 million in missing box office dollars went. (Spoiler: straight-to-streaming.) All that, plus new episodes of The Secret Lives of Mormon Wives, Supergirl heads to streaming, a new YA show does well on Netflix, the viewership for the Daily Wire’s $50 million fantasy show, what TV show hit-maker can’t make streaming hits, CD sales, all the flops, bombs and misses, and a whole lot more. Let’s dive right in! (Reminder: The streaming ratings report focuses on the U.S. market and compiles data from Nielsen’s weekly top ten viewership ranks, Luminate’s Top Ten Data, JustWatch and Reelgood interest data, Samba TV household viewership, company datecdotes, Netflix hours viewed data, Google Trends, and IMDb to determine the most popular content. While most data points are current, Nielsen’s data covers the weeks of September 7th to September 13th, 2026. You can find a link to my terminology here.) Subscribe Television - Ride or Die Is a Hit…But It Got Cancelled? What The Data “Says” I hate the phrase “here’s what the data says”, even though sometimes I find myself using it. If data just told us everything, we data folks would be out of a job! Data is often messy, complicated and nuanced. When it comes to strategy, many of the best decisions can’t actually use “data”, since it’s vastly too complicated to model. Anyways, I’ve been singing the praises of Prime Video’s Ride Or Die this summer, a show that did very well for them. And yet…it got cancelled. Now, usually when a show gets cancelled, a bunch of websites say “a hit show got cancelled”... Here’s the thing, though: in this case, they’re right! The data “says” that this show was a hit. Specifically, out of 522 first seasons to make a week on the Nielsen charts since 2020, it’s the 75th biggest. On Amazon, Ride or Die is eighth out of 41 first seasons: That’s top 15th percentile all time, almost exactly, which is my definition of a “hit”! I mean, outside of Reacher, Amazon needed a hit this summer! (Admittedly, Off Campus did well for Amazon in the spring, but Ride or Die performed even better.) So what happened? Why did Amazon cancel one of their rare 2026 bright spots? Well, Deadline reported a reason—which I’ll get to—but I’d break it down into three relevant questions: What is a given streamer’s reasoning? Is that reasoning sound? Could the stated reason not be the real reason? According to Deadline, Amazon felt the show “over-indexed” in middle-aged women. (This was based on leaks from within Amazon.) The logic is this: Amazon already reaches this demographic because of unlimited two-day shipping bundled with Prime, so they don’t really need more shows like this. Now, it’s my job to ask if that’s a good reason, and I’ll be honest, in this case, I don’t think it is. Sure, this show may “over-index” in middle-aged women—I’ll just assume that’s the case here—but I think folks often oversell how important over-indexing actually is. When a show is a “hit”, it isn’t a bit bigger than its rivals; it’s usually multiples bigger. For example, Elle only made the charts for two weeks at 8.3 million hours each. So Ride or Die was 80% bigger in its first two weeks, and likely had a much stronger hold, with its terrific 13 million hours in week three. Ride or Die was nearly three times bigger than Sterling Point through three weeks too. And that means that _Elle, Sterling Point a_nd other Amazon YA shows need to not just over-index a little with younger women (assuming that’s who Amazon prefers over middle-aged women), but massively over-index. Otherwise, more people in that demo (and a bunch more besides) likely watched Ride or Die. (I’ll try to explore this concept more in a future article.) And don’t get me started on how Ride or Die did compared to a certain (very, very popular, possibly the world’s most popular) YouTuber’s reality show…which didn’t make the charts after its first week and likely cost much, much more than Ride or Die. All to say, Amazon-MGM Studios/Prime Video can tell reporters that this show didn’t reach the right audience, but that excuse feels weak to me. So this leads to the third question: could other reasons have come into play? Yes! Renewal or cancellation decisions rarely (I’m tempted to say “never”) boil down to one variable. At its simplest, it’s two things: budget versus viewership. But often factors like critical acclaim, ownership and, yes, personal opinion come into play. In this case, Amazon has a new executive running things, and this show isn’t owned by Amazon-MGM Studios (unlike Elle). And yes, it may not have indexed with the right target viewers, too. Likely all of those factors came into play.1 My guess is ownership ended up mattering most. Amazon doesn’t own Ride or Die. It’s actually produced by another major studio. Fortunately/allegedly, this show is being shopped around. I can think of one brand new CEO who should consider it for one of his two major streamers…especially if he wants to rebuild goodwill in Hollywood. His name is on this list of exec producers of Ride or Die… If David Ellison needs some goodwill—and he does—rescue this female-led show tomorrow and grab some good headlines for a day. The data justifies it. Quick Notes on TV We’re just getting started with this issue, but the rest is for paid subscribers of the Entertainment Strategy Guy, so if you’d like to find out… How The Secret Lives of Mormon Wives latest season did on Hulu, post-controversy... Whether Supergirl soared on HBO Max... Why the box office is losing money due to the streamers... The viewership for the Daily Wire’s $50 million fantasy show... Updates on The Gentlemen, Lanterns, Outer Banks, King of the Hill and more... What former hit-maker has three streaming flops in a row... Whether The Mandalorian and Grogu popped in week two... All the flops, bombs and misses... And a whole lot more... ...please subscribe! We can only keep doing this great work with your support. If you want an idea of just how much content you’d get in a full issue, check out this older issue. Coming Soon! Oh man, we’re almost caught up! We just have two more issues, then the streaming Ratings Report will be coming out just two weeks after a TV show premieres! Next issue, I’ll be looking at the first two weeks of the NFL Thursday night games on streaming, a huge new spinoff, Reacher’s Neagly on Prime Video, and the latest edition of Monster: The Lizzie Borden Story on Netflix, the return of Slow Horses on Apple and MobLand (quietly one of their biggest shows) on Paramount+, and Dancing with the Stars on both ABC and Disney+. Plus I’ll take a look at animation for adults in 2026. The week after, a ton of movies are coming to streaming: Toy Story 5 on Disney+, Backrooms on HBO Max, Jackass: Best and Last on Paramount+ and The Breadwinner on Netflix, plus a Unabomber film on Netflix and a romcom on Prime Video. Woody Harrelson and Matthew McConaughey have a show on Apple, along with a game show inspired by Willy Wonka that everyone seems to hate. Long term, some crazy (but hopefully crazy like a fox) IP news. FX ordered a new Sons of Anarchy show from Charlie Hunnam, but it’s not about the biker gang. No, it features the show’s cast playing themselves in a new “meta-thriller”. Huh. And a lot of the cast is coming back. Next, Ryan Gosling is producing a Flintstones movie about a “grown-up Bamm-Bamm Rubble”. Also, huh. This one is in “early development” so we’ll see what happens. Both projects are based on IP, but also seem to have crazy takes on the IP, which I like.
September Monthly Review: Health MCP Guy?
For all the time I spend talking about APIs and data exchange, this newsletter has had embarrassingly little interoperability of its own. It’s old-school SaaS, baby - UI or bust. Not all of that is my fault! Substack, for all its popularity, is (as mentioned before) really frustrating as my system of record. The primary user it optimizes towards is the reader, whose attention feeds the app, the Notes feed, and the recommendation network that Substack is actually selling. Health API Guy is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe They’re falling into the same trap that Medium once sprinted into, albeit at a much more glacial speed. Younger generations may not remember (or never knew) it, but Medium once lured publications like The Ringer and Signal v. Noise onto its platform, only to lose them as it rebuilt itself around a paywall and an algorithmic feed. Substack’s original pitch circa 2020 was the antidote (you own your list), but every new feed feature chips away at it. I get it: writers are generic supply in the eyes of the Substack product team. Generic supply follows demand, so the consistent pressure on the PMs is to generate demand, which leads to every product cycle going toward discovery (Notes, recommendations, the app, live video). The end result lackluster author tools (laggy and broken analytics are truly crippling) all despite the premium price point. The archive is where it crosses from authorial gripes to actual reader impact, though. Across Substack and LinkedIn, I’ve published 1,218 posts and nearly 900,000 words since 2018. Substack’s native search can see only the 464 posts that live on Substack, and even there it’s woefully inadequate at surfacing relevant content (even for the guy who wrote it)! Readers have the same problem, writ large. Ask the Archive Fixing search (for both my readers and my own use) turned out to be extremely non-trivial, unfortunately: There’s no public API for readers to access content RSS exists for new content notification, but it only carries the 20 most recent posts and cuts paid ones off at the preview The only Substack MCP is an admin tool for analytics The export exists for authors, but it’s painfully awkward and slow. So I built something to work around that: the Health API Guy MCP. It’s a private MCP server holding the full text of everything I’ve written across Substack and LinkedIn. Add it to Claude or ChatGPT as a connector and you can ask things like: It pairs keyword search (for acronyms and case names, of which we have many) with semantic search. Every post is also labeled by topic, form, and court case, so you can pull a lawsuit’s (or multiple cases’) entire docket across Substack and LinkedIn in one query: Hail, Healthapiguyhydra This is also a bit of me eating my own cooking. I’ve spent the year arguing that headless healthcare is arriving and that developers want primitives over abstractions. Last week in Hail, Benihydra, I watched Salesforce plug its platform into Claude and bet that the records underneath outlive whatever interface sits on top: Salesforce must believe that the Hydra’s immortal head is the records, permissions, and business logic (in that they will survive whatever happens to the interface). It felt a bit hypocritical to keep my own archive behind a Substack search bar. The archive is the body. Substack is one head, and now whatever tool you might use is another. Will the EHRs follow my lead? Probably not imminently, but one can hope! Why Founding Members Only The MCP is available to Founding Member subscribers. A few reasons: It costs money and time to run. Hosting, the database, Clerk usage and AI compute for search and labeling are significant operational costs (that increase with usage). Substack gives me no way to automate access. There’s no API to check whether someone is a paid subscriber, so every user is manual operations for now. It’s a beta. We’ll need to improve things and parts will break. I’d rather they break for a small group of dedicated readers who will tell me about it. The Founding tier deserved a real benefit. Until now it was mostly a generous tip. Now it comes with something. If you’re already a Founding Member, reply to this email with the address you’d like provisioned. If you’re not and want in, upgrade here. Larger teams (such as companies) can get access too: group subscriptions at the Founding tier come with the group discount. The admin for the subscription can reply with the list of addresses to provision. Get 20% off a group subscription Pricing Founding membership is now a flat $200 a year, replacing the old suggested $150 that readers could adjust. Existing Founding Members will have access grandfathered in at their previous rate. Month in Review Articles Published Make Networks Dumb Again (Sep 11): AI is making old analog networks cheaper to use and flooding them with machine traffic in the process. A plea for agent-native rails with the intelligence at the endpoints, and for Washington to drag them to ubiquity. Interoperability, Unbuttoned (Sep 29): Interoperability is a series of workflows in a trenchcoat, and this post finally takes the trenchcoat off. A color-coded map shows which provider workflows went digital, which are fragmented, and which still run on fax. Video Content Antitrust’s System of Record Problem (Sep 3): The FTC’s investigation into Epic prompted a ninety-minute webinar, now posted with chapters and a smoothed transcript. We get into why market definition, missing B2B share data, and the SSNIP test leave antitrust poorly suited to systems of record, plus why information blocking may be doing more than any pending lawsuit. The Information Exchange: Cerner Milking Edition (Sep 18): This was so fucking fun. Ryan Tucker, Ryan Brickner, Brad Thorson, and I try an Around the Horn format, complete with points, penalties, and a mute button. We debate whether Oracle Health is dead, who would actually buy Cerner (IBM enters the chat), p(Doom), and why the front desk still hands you a clipboard. Regulatory Ask Not What Your Network Can Do (Sep 23): CDC’s RFI on public health data intermediaries lists use cases already served, unevenly, by eCR, lab reporting, IZ Gateway, and syndromic surveillance. The real opportunity sits in one question about AI: a network that can handle the next emergency without another interface project. Court cases Particle v. Epic: The Naughty Schoolteacher (Sep 4): After a summer-long vigil, the answer on Epic’s early summary judgment bid is two paragraphs and a remedial assignment. Particle advances to the next grade, and market definition gets another year of discovery. The Epic Discovery Files (Sep 8): Texas’s motions to compel arrive with a 251-page appendix of requests, objections, and meet-and-confer letters with Cravath. Inside: Epic using generative AI for document review, denied interoperability requests in scope, and Noerr-Pennington raised against lobbying discovery. Epic v. Health Gorilla: “Working as Designed” (Sep 9): Epic wants nothing to do with the data breach MDL, and explaining why puts fresh discovery on the public docket. The spiciest allegation: Health Gorilla told GuardDog to scrub its law firm business from its website instead of cutting off access. Veeva v. Epic: Legal Bypass Surgery (Sep 15): Veeva asks the Wisconsin Supreme Court to skip the Court of Appeals, while law professors and three policy groups line up as amici. Meanwhile, Epic’s lawyers apologize for another round of citation errors, including a quote that never existed. Texas v. Epic: Behind the Black Bars (Sep 17): The court portal served up Epic’s unredacted filings, revealing a $5 million document review budget and the AI workflow behind it. Also inside: the 200-plus company names Texas wants searched, and one very expensive Zoom invitation. OpenEvidence v. Doximity: Consequences Will Follow (Sep 25): OpenEvidence tells the court it will not comply with a discovery order, and its own lawyers at Quinn Emanuel say they won’t defend the choice. Add years of Slack reported unrecoverable, and Rule 37 sanctions come into view. Health Tech Keeps Choosing Violence (Sep 28): Doximity moves to end OpenEvidence’s case, Epic asks Texas to show its work, and a sleepy Audacious Inquiry patent fight sprouts antitrust and information blocking counterclaims. The main event is a long-sealed Cognizant v. Infosys ruling that hints where system of record antitrust claims can survive. EHRs One Verb for Oracle Health (Sep 14): Oracle’s earnings call quantifies every data center to the decimal point, while Oracle Health gets a single verb: “accelerate.” Panning the few healthcare mentions turns up an agentic care management system, a research-to-care ledger play, and the question of whether healthcare gets Larry’s patience. Epic’s Forretress Under Siege (Sep 30): Judy Faulkner tells a panel Epic paused hundreds of projects to harden its software and calls it a “shame,” while a spokesperson insists the roadmap hasn’t changed. Project Glasswing, a Pokémon-branded QAN sprint, and a trove of Jodel grievances reconcile the two Epics and preview the Jevons paradox coming for all software. Industry Analysis Scanning for a Network (Sep 1): Scan.com raises $220M to build the imaging network labs never got, speedrunning Zocdoc’s consumer-to-B2B pivot. Its wedge to network density may be the most American one available: personal injury litigation. The Front Desk Wants the Damn Whole Building (Sep 16): Hello Patient buys Converse Health, pushing its AI front desk into referrals, chart work, and prior auth. The copilot categories keep blurring, and your distribution partner is still your final boss. Cross-industry Comparisons A Different Patchwork Quilt (Sep 10): Healthcare loves to envy open banking, but the US version has no mandate, a toll booth at Chase, and data quality problems of its own. The grass isn’t greener; it’s another American patchwork quilt, and perhaps a worse one. Hail, Benihydra (Sep 21): Salesforce brings its platform into Claude, Slack, and its own agent, inviting the interfaces that might replace it. Cut off one head and more grow back, so long as the records underneath stay put. Living on A Trade Secret Prayer (Sep 22): Six credit unions are suing Fiserv, the Epic of core banking, by claiming their own member records as trade secrets. Banking has no Cures Act, so its lawyers reach for legal alchemy that healthcare has already tried. Other News None this month External Media Nabla Accelerate NYC: I kicked off fall conference season at Nabla’s Accelerate event in New York, where Delphine Groll opened by asking the room to picture the company it wants to become and the one it doesn’t. I liked her framework of bureaucracy, fear, and ego as the blockers. I also love vendor conferences/events in general - different vibe than industry biggies. Healthcare 101 with Nikhil Krishnan: I briefly guest lectured Out-Of-Pocket’s Healthcare 101 class. Appreciated Nikhil’s takeaways cover whether AI doctors can query HIEs, why healthcare’s data-liberation rules beat most industries’, and the hi-res graphic is in the comments. HCN: Epic Probed by the FTC: I joined the HCN guys to talk through the FTC’s investigation into Epic. Cedric’s new hairdo was an unexpected plot twist. The episode also covers MFN drug pricing, healthcare M&A and Eli Lilly’s acquisition streak. Fall Conferences: At eHealth Exchange’s Annual Meeting in Austin on Oct 27, I’ll moderate a panel on the tensions between AI, HIEs, and timely patient access with Jean Ross (Primary Record), Therasa Bell (Kno2), Byron Crowe (Doctronic), and Michael Marchant (Freenome). Here’s the rest of my schedule if you want to meet up: CommonWell (Redwood Shores, CA): Oct 13-14 Will be participating in a fun Jeopardy event Open@Epic (Madison, WI): Oct 20-23 Will be enjoying Verona, Willy Street, and other local adventures eHealth Exchange and Sequoia Project (Austin, TX): Oct 27-29 Yeehaw HLTH (Vegas): Nov 15-18 Groundhog Day, Venetian Edition RSNA (Chicago): Nov 30-Dec 2 See me at “Governing AI in Breast Imaging: What Health Systems Are Actually Doing” Digital Health Counsel 2026 AI Summit (Seattle): December 2-3 Will likely be speaking to a bunch of lawyers about information blocking Posts I Liked ChatGPT’s Epic integration: game changer or nothing burger?: Joshua Liu, MD offers five thoughts on OpenAI’s Epic launch, starting from the point that it rides the same open APIs available to anyone. He argues read-only access leaves ordering to Art and the ambient startups, and asks why HCA is piloting it alongside Commure. He’s sorta a must-follow at this point - consistently excellent, thought-provoking takes. Epic, Lilly, and the Discovery pass: Seth Chaney reads Epic’s first major Discovery deal (with Eli Lilly), as well as its revised developer terms on commercial displays, as the pipe finally opening between EHRs and pharma. Life sciences and health continue to collide. Scott Rossignol on EHR payer platforms and ePA fees: Scott argues the top ambulatory EHR vendors are turning the prior auth mandate into a per-member-per-year toll on payers, for an API whose operating cost has nothing to do with covered lives. His fix is two lines of regulatory text: put the CRD/DTR/PAS certification criteria in the Base EHR definition, and bring payer-facing ePA connectivity under § 170.404’s cost-based fee rules. It’s a good solution. Health API Guy is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe
The Big Holes in Wearable Heart Rate Variability And Readiness Scores
On September 9th, Apple announced it was revamping its Apple Watch Health Sensing System, rolling out a Readiness score (0 to 10), and increasing the frequency of heart rate variability (HRV) outputs 24-fold. This can be viewed as upping its competition with various other consumer wearable sensors. These “readiness scores” as a composite of multiple metrics, with heart rate variability (HRV) being front and center for most. The majority of Americans are now using wearable sensors, which equates to well over 100 million adults. HRV and Readiness scores are increasingly being marketed as a measurement of autonomic nervous system health, a digital marker for future disease, a clock for biological age, and a holistic metric to promote healthspan and even longevity. (Apple also introduced a new longevity tab and “Health Age”.) None of this has been proven. In this edition of Ground Truths I am going to review what we know about heart rate variability and readiness scores. Heart Rate Variability HRV is the variation in normal heart cycle timing. The variability of the heart rate, the barely perceptible millisecond changes in time between consecutive heart beats (see R-R intervals in the Figure below, left panel), is due to interplay between the sympathetic and parasympathetic (vagal nerve) inputs. Distinct from heart rate, individuals with the same heart rate can have very different HRVs. It is a rough reflection of the autonomic nervous system (ANS) activity, inadequate to say whether a person’s ANS function is abnormal. For more than three decades, heart rate variability (HRV) has been measured and several studies have found an association of low HRV and clinical outcomes, particularly a link with higher all-cause and cardiovascular mortality. There have also been less well established links of low HRV to risk of early cognitive impairment, dementia, mental illness, Type 2 diabetes and substance abuse. An important reminder is that HRV is a surrogate marker without any established cause-and-effect relationship. If you increase your HRV, that doesn’t mean it will improve health outcomes. In fact, there is no hard evidence for that. All that work linking to health outcomes was done with electrocardiogram (ECG) derived HRV. Now, in the era of consumer wearables, this is getting assessed differently, by optical pulse (yes, the lights you see) plethysmography (PPG) or what is called pulse rate variability (PRV). They are not the same, as shown below (right panel) and only concordant when the delay between the ECG and pulse is kept constant, which basically means at rest. I should mention there’s also what I will call MPV, a mattress mechanical movement sensor, a derived heart rate variability, that companies like Eight Sleep use, even further away from directly measuring HRV. HRV has not one uniform measurement but many different types of quantification, such as RMSSD, the magnitude of difference between successive R-R intervals of normal sinus beats (N-N) or SDNN, the standard deviation of NN intervals, both in milliseconds. SDNN is one of the so-called frequency domain HRVs (others are LF, HF, LF/HF). Different wearable sensors use different metics; Apple has relied on SDNN and nearly all of the others use RMSSD, which is generally considered the more accurate metric. There’s also the different length of time measured, such as for a matter of minutes, all day, or an overnight’s sleep. Short measurements are especially problematic since they don’t capture enough of respiratory modulation and other factors that influence HRV. Share Ground Truths How well does HRV correlate with PRV? There are very limited studies, especially independently done. One that is commonly cited was conducted by Air Force researchers in only 13 healthy adults assessing Oura ring 3 and 4, Whoop 4.0, and Garmin Fenix 6 and showed a correlation coefficient of 0.88 to 0.97 and a mean absolute percentage error from 6 to 10%. The correlation is not a perfect 1.0, but there’s at least a fairly high level of correlation. HRV is supposed to increase during the night due to takeover of the parasympathetic nervous system, and higher during deep sleep. A recent example of my 1 week, all day “HRV,” and one during sleep is shown below. As you can see, the N of 1 data are inconsistent for the same days from different sensors (Oura, AppleWatch, Fitbit Air, and Eight Sleep) by patterns, absolute numbers, and comparison with prior days and weeks. The largest study in over 8 million Fitbit users (the old version, not Google Fitbit Air, introduced in May 2026) gives you a sense of the effect of age, sex, and the 2 different main HRV (here PRV) metrics, with RMSSD on the left and SDRR (=SDNN) on the right below. That study, from data collected in 2018, is a major outlier, since all the more recent ones are tiny with respect to sample size. Many of the companies have not had independent evaluation of their HRV, such as Eight Sleep, but have published a low standard error on their website. There are some other published studies on the correlation between HRV and PRV, but they are all small and only in healthy adults. A scoping review emphasized the lack of study in underrepresented individuals, including the aged, people of color (which affects the PPG signal), and individuals who are underweight or obese. Add the typical adult age 60 plus with one or more chronic diseases. For example, one study in over 900 adults found poor correlation of HRV and PRV, non-uniformly underestimated across many chronic diseases (cardiovascular, endocrine, neurological, respiratory, and others), concluding PRV is “an invalid surrogate for HRV.” A recent systematic review of 43 studies comparing HRV and PRV found reasonable pooled absolute standardized error (HRV as gold standard) but only 10 of the studies provided quantitative synthesis in ideal conditions. Their main conclusion was similarly cautious: “PPG-derived HRV [PRV] should not be regarded as universally interchangeable with ECG-derived HRV across all devices, populations, and recording contexts.” Factors Affecting HRV and PRV That gets me to the long list of factors that affect HRV (and PRV) besides the device, the type of measurement (RMSSD, SDNN or others), the person’s signal, the sensor site, the duration of data capture, if weighting by sleep stage is used, how artifact is processed and corrected. And this list is not complete!: Oura puts out data from their community of users (who input data) on what affects their overnight HRV. The factors currently provided are: no alcohol (increase 8%), melatonin (increase 2%, float tank (increase 2%), wine (decrease 4%), and party (decrease 14%) in overnight HRV. Must be some big parties! Subscribe What is a PRV measurement good for? It has been falsely characterized as an index of “autonomic balance” and a specific indicator of stress. A 2018 review of the studies available for HRV and its relationship to stress, not using any of the current wearables, found that stress can lower HRV. But so can many other factors. The non-specificity of the signal, indexed to the table above, is striking. Evidence from a UK Biobank study of over 46,000 participants with actual HRV looked at genetically predicted HRV, a genetic risk score, that failed to show the expected HRV-mortality link, indicating that _HRV is likely not causa_l, but rather a reflection of person’s physiologic state. A review of consumer wearable HRV data from 5 longitudinal studies showed that nighttime PRV was not associated with perceived stress, and surprisingly higher HRV, in the largest cohort (N=717 participants), was correlated with higher stress. An Oura ring cohort of 525 first-year college students found a link between overnight PRV and perceived stress, but that was also seen with resting heart rate, sleep, and respiratory rate. Several very small studies have examined the relationship of HRV and athletic injuries or guiding training with mixed, and predominantly negative results. HRV biofeedback training with paced breathing had no significant effect on reducing stress or raising HRV, as demonstrated with sham controlled trials. When HRV for multiple days showed a decline in conjunction with body temperature, the Oura ring published data for prediction of Covid. The WHOOP company sponsored an observational study, published in 2026, of 30,000 users for 72 weeks, without a control group, that reported reduced alcohol intake (5.8 % points) by self-report. That doesn’t tell us much, and particularly about the merits of HRV for behavioral change. If you use the same device and conditions as longitudinal trends for multiple (at least 2-3) weeks that may be the one way to get something useful from the measurements. Data for overnight sleep with minimal motion and using RMSSD is the best proxy for real HRV. The reason to look at trends rather than any given night is that it more likely represents something, even though you won’t know with certainty what the “it” is. Keep in mind there are no data, no peer-reviewed evidence, to show that HRV fluctuation in-person has any correlation with health outcomes. Share Readiness Scores These are proprietary scores that integrate different metrics for each of the wearables: no algorithms have been disclosed. They are unvalidated against health outcomes. In a review of 14 composite health scores of readiness and recovery, HRV contributed 86% to the scores, followed by reading heart rate (79%), physical activity and sleep duration (both at 71%). That review noted the substantial variability in measurement protools and lack of standardization. Sleep staging is notoriously inconsistent and inaccurate by these sensors, which adds further to the HRV uncertainties for what the scores, which use sleep stage data, mean. Only resting heart rate has been shown consistently across devices to be extremely accurate. I’ve made a Table to summarize what we know about which metrics are included, the scores, any peer-reviewed studies that compared the readiness score with health outcomes, and the corresponding (if any, NA-not available) citation. You will note that some companies do not use the term “readiness," such as WHOOP for recovery, and Garmin, which has 2 different scores, one of which is Body Battery. Eight Sleep uses the term “Fitness Score.” They all include HRV; Apple includes a new metric they call “Recovery HRV” which among other components uses 7-days of sleep, but it is unclear what this means or how it is differentiated from other scores (there are clearly no data for outcomes). We have no knowledge of how the different components are weighted or whether any of these scores are better than resting heart rate, HRV alone, physical activity, or any other single metric. Since none of these are standardized, they are not interchangeable, so if you get a 90 for Oura that has no relationship to a 90 on a Google Fitbit Air. Notably, the company can update its algorithm for readiness score at any point without notification to device users. Without any useful evidence of actionability for these scores or established relationship with health outcomes, it is hard to make a case for their value. At the Apple recent announcement they showed their Readiness score (0-10) on the watch (Figure below) but there are no published data on this score, not even on their website. It’s available only on their new Watch Series 12 or Ultra 4 [of course, ;-)]. That exemplifies the problems with these scores, lack of data and evidence for being meaningful to promote health. Perhaps the best study (which isn’t saying much) is the WHOOP Recovery for golfer performance, because it did correlate with an objective outcome, even though there was no control group and the authors were all from the company. Among the 389 pro golfers, an absolute 10-per cent point increase in Recovery score was associated with about 0.5 fewer strokes per round. But that’s hardly a health outcome! WHOOP is also conducting a study in over 2,700 runners to see if their recovery score will be linked to less injuries and improved performance, but that is not yet published and has no control group or randomization. Putting This in Context For two decades I’ve been enthusiastic about the potential for digital health and particularly wearable biosensors. Over the years, we’ve seen some great progress for their ability to promote physical activity and accurately detect atrial fibrillation (the first FDA cleared deep learning AI for consumers). That work was the subject of rigorous research. But there are holes in the data and evidence for other metrics. One notable one is the “VO2 max” story that I wrote about earlier this year. At that time many subscribers asked me to cover heart rate variability, which I finally got to here. When I dived into the research and publication for HRV and readiness scores, I expected to find at least some that were of high quality and demonstrated their utility by linkage to health outcomes. To my surprise, I found none. The wearable sensor measurements for HRV (PRV) are, for the most part, accurate, but that validation work has only been done in small studies of healthy adults and does not take into account the long list of factors, from the device, software side, and the user side, that affect HRV measurements. Moreover, this metric chiefly relies on optical sensing and, as we have learned for heart rate PPG sensing, may be less accurate in people of color. Keep in mind that all of the health outcome association evidence comes from ECG-derived HRV; none are from wearable sensor data. I will repeat the key point: there's no peer-reviewed evidence to show that in-person HRV fluctuation—or efforts to raise your HRV— has any correlation with health outcomes. For those of you who look at your HRV on awakening, or even 2+ week trends of it being low, I hope this context helps to relieve any anxiety. Yes, low HRV (not PRV) has been shown to increase risk of some diseases as summarized above. But efforts to raise your HRV—a surrogate metric— has not been established for improving any health outcomes. HRV does not have any evidence of causality (the genetic evidence actually goes against this possibility). In this summary, I have not included data for other wearable sensors such as Polar, Samsung, Withings, Suunto, Amazfit, Coros, Ultrahuman, or additional mattress sensors. These are beyond my first-hand experience, and as far as I know from my in-depth review none have any peer-reviewed published data that differ from the 6 sensors I’ve reviewed here. From the points I’ve gone over above, we’re not ready for readiness scores. Besides being proprietary, they are predominantly based on metrics that have their own issues. It’s compounding the problem, like building a house without a solid foundation that has never undergone a rigorous inspection, and then selling it. Like I mentioned for PRV, you can look at trends over weeks rather than any single day, to get a handle, but even that may not be helpful. There’s simply no evidence that these scores meaningfully relate to health outcomes. I’d emphasize they might, but that requires doing prospective or randomized studies to prove it. None exist. There’s great promise for HRV/PRV utility_. For example**,**_ Prof Maiken Nedergaard, who discovered the brain glymphatics that are essential in eliminating metabolic waste products from the brain during sleep, has posited that HRV could be a non-invasive marker for neuromodulator oscillations, brain-body regulatory circuits, and brain clearance. That would be extremely useful, but like everything else on HRV and readiness scores it requires solid research and validation. The lay media isn’t helping much to get the story straight. Earlier this year The Economist published a piece entitled “The most useful indicator of your overall health” which ordained HRV as an “accumulated stress score.” That’s akin to the false assertion about VO2max: “V02 max is the singular most powerful marker for longevity.” As I’ve summarized here, that is not established. The fact is that so many things can lower HRV, including physical exercise (especially an intense workout), reduced sleep quality, stress, the list above, no less the device, signal, and software. Whatever fluctuations observed have not been correlated with any health outcome. Sadly, “datamaxxers” are widely using HRV and readiness scores that have never been validated to mean anything. We already know that for some people using the sensors for sleep metrics, it can induce “orthosomnia,” an obsession to get high sleep quality, with associated high levels of anxiety. In an experiment done by a company to promote sleep quality for its employees, “For those employees who did use the trackers, many reported feeling perfectly rested until their tracker told them they had had a terrible night. Others were told that they had slept like a baby when they had actually been lying awake worrying about the quality of their sleep. “ The same problem can result from preoccupation with HRV or readiness scores, with anxiety that would lead to further reduction in both. It you are using a wearable like >100 million American adults, it’s OK to look at these data, but contextualized with the major caveats reviewed here. If you are one to require evidence that HRV or readiness scores are linked to health outcomes, you may not even want to look. The companies make it hard to turn them off! Let me end with the companies that make and sell wearables. Apple’s doubling down on HRV (24-fold more reporting and heart rate very 5 seconds) and introduction of a Readiness score tells us that consumers have bought into these metrics and they are joining the club. However, all of this is occurring with a backdrop of tens millions of users, claims about the data that are not backed up by adequate evidence, marketing way out in front of whatever limited data exists, and not being transparent about their readiness score algorithms. The companies can well afford to do the research that is needed to connect these metrics with health outcomes show, once and for all, that increasing HRV or using readiness scores promotes our health. If they believed and invested in the products they are selling, we’d not be in this position of not knowing. That’s essentially where we are with HRV and readiness scores. Perhaps someday this will change and we’ll have good reason to embrace them. NB: I wrote this post. No AI. I have no conflicts of interest with any of its content. Loading... Ground Truths has 215,000 subscribers from every US state and 214 countries. There are over 300,000 followers of Ground Truths so more than 90,000 folks who can easily convert to be free subscribers. Your subscription to these free essays and podcasts makes my work in putting them together worthwhile. If you’re not a subscriber, please join! If you found this interesting PLEASE share it! Share Ground Truths The proceeds from all voluntary paid subscriptions go to support our summer internship program. It enabled us to accept and support a record number of 62 summer interns that joined us in 2026! These are high school, college and medical students selected from thousands of applicants. We couldn’t do this expanded program without the funds coming in through Ground Truths. Thank you!
15 discouraging hours with Onimusha: Way of the Sword, and 10 brilliant ones with He Who Watches
In the last few weeks, while playing some video games that I might review, I emailed two cries for help. I sent the first to a rep for Capcom, confessing that the company’s samurai game Onimusha Way of the Sword (out today for PC and console) was crushing me. “Feels like I’m playing ‘hard’ mode, not ‘normal,’” I wrote. “And I don’t want to have to knock it down to “story”. I’m exploring, trying to level my gear, but these two-phase boss fights with no mid-fight checkpoints are brutal.” I wagered I was only about halfway through the game, if that. I wasn’t sure I could keep up with the steepening difficulty curve. Was I missing something? Then, on Monday, I took some photos of some puzzles that were baffling me in the mind-bending He Who Watches (out now on PC), wrote out lengthy descriptions of how I’d been trying to solve them and asked Bobby Vanden, the game’s solo developer, for a nudge in the right direction. I told Vanden I’d spent hours on these puzzles. “I’m stuck in triplicate,” I told him. Could he tell me if the solutions I was attempting were even on the right track? I didn’t mention that the puzzles—coupled with major time pressures from other work—were stressing me out so much that playing the game was giving me headaches. This double dose of panic was unusual. I’m used to managing a workload full of reporting and reviewing, used to a crowded schedule and the need to stay calm amid the bustle. But I was stressed about these games. No one was forcing me to review them, nor even play them. I’d simply found them interesting enough to start, hoped they’d be interesting enough to write about and I’d sunk hours into them assuming I’d made the right call. Suddenly, however, I felt trapped. I was emailing for a way out. Above: This guy took me about an hour to beat. The feedback I got from Capcom confirmed my fears. They didn’t say it, but I will: Onimusha Way of the Sword is not for me. It’s a brutally hard game in need of a normal difficulty level. It’s a game that spikes so severely during its abundant boss fights—yet offers such flawed ways to improve when outside of them—that it feels like a bait-and-switch. In the new Onimusha, you play as famed samurai Musashi Miyamoto, aided by a magical soul-absorbing gauntlet that is possessed by a lady who offers advice. (Way of the Sword is a continuation of Capcom’s long-dormant Onimusha series, but that set-up had me wondering if the company was cribbing from their ill-fated Bionic Commando reboot; read about “wife-arm” when you get a chance). Way of the Sword is partially set in 17th century Kyoto, an explorable city where Miyamoto can fight demons, help civilians and eventually rescue ghostly dogs. The city evolves over time, with new mini adventures and obtainable loot on offer, all feeding into systems that promise to upgrade your gear and sword-fighting skills. Upgrades, however, are so incremental as to be imperceptible, causing the exploration of Kyoto to feel like inconsequential busywork. From that hub, players can go on linear missions into forests and old temples, where they will encounter many of the game’s tough bosses. Combat in Way of the Sword is akin to many modern melee games. Players are asked to mix attacks with dodges and parries. They’re encouraged to deplete an enemy’s stamina bar in order to do greater damage to its health bar. The old Onimusha games’ signature move was the issen, a counter-attack so sudden it’s all but pre-emptive. You strike just as they begin to strike toward you, beating them to the cut. By definition, the issen is a high risk move, leaving you wide open to pain. It feels risky in the game, especially against bosses, who can maul you in just a few blows. A game might train you to perform an issen well by offering a variety of lesser enemies who are not as tough as the boss but get close. Or it might offer a toehold against bosses by saving or checkpointing the player’s progress when a boss’ first health bar has been destroyed and before the second health bar appears. It might offer the less skilled player a chance to over-level their character, should the player simply commit to playing through its optional sidequests. Way of the Sword does none of this. Fighting its trivially easy grunt enemies and then its ferociously adept bosses is like playing a game of chess where I’ve got a rook, while the opposing side is all pawns and five queens. Bosses simply don’t checkpoint. You may be able to clear a boss’ first phase without taking much damage. Good for you. Do it again, every time you fail the second phase. As for the the central Kyoto hub, I kept pace with all of its added side activities and never felt like I was overpowered. Bosses could still easily buckle me. The new Onimusha, I concluded, is just too discouraging to keep playing. The advice I got from Capcom had been well-intentioned. They reminded me of the game’s various systems. They alerted me to an optional assist that flashes recommended button inputs (basically: counter this move; dodge that move). It was a little helpful. So was dropping down to “story mode,” which I tried last night. Story mode turns on those button prompts by default. It makes countering and dodging way too easy, nearly automatic. I don’t consider myself a master of action games like these, but I’ve cleared the likes of God of War, Stellar Blade, Star Wars Jedi and others of sub-Sekiro difficulty and had a fun time doing so. Not here. Not yet. Unless they patch Onimusha: Way of the Sword and add a middle difficulty or just checkpoint the damn bosses when they’re between health bars, my streak of loving Capcom games released in 2026 will likely end at two. It’s a shame, because I was having a good time finding those ghostly dogs. And then there’s He Who Watches, a phenomenal first person puzzle game with Portal vibes, except instead of a portal gun you’ve got a magical bow that shoots one arrow that can activate switches or be used to drag blocks. The twist is an actual twist: you can also walk on walls and the ceiling, reorienting gravity as you do, thereby placing blocks onto those other surfaces. Walking on walls and ceilings to solve problems is really the best thing. Would that I could in reality. He Who Watches consists of locked-room challenges, each brief enough to only require about a minute to actually solving, though it will take five, 10, eventually 30 minutes to figure out said solutions. They are not obvious! It’s a game of layered mechanics and shocking solutions. For example, see if you can follow this: You can drag a block to a wall. Then you can walk on the wall, which now feels like the floor, and drag that block across the wall. Later, back on the floor, you might realize you can look up. The block is still on the wall, stuck to it. You can shoot your magic arrow to its underside. Doing so creates a tether, so that you can yank the block off the wall and carry it, as if it were a balloon, into the spot where you need it to be to solve the puzzle. Here, watch Vanden explain some of the game’s basics in this launch trailer: As with any great puzzle game, the solutions to He Who Watches’ puzzles elicit relief and amazement. Oh my god, that worked! I didn’t even think that could be possible! Learning to solve the game’s puzzles feels like learning magic. But, oh, does the game get hard. Its Steam reviews, mostly glowing, are full of players remarking how quickly He Who Watches made them sweat. Wrote one: “loving this one already although my brain is hurting!” Wrote another: “As you get further in the game it feels like you need an IQ level of at least 120+.” Vanden has shown some mercy. Each locked room puzzle includes an optional junior version of the puzzle, meant to teach you. Difficulty nevertheless ramps severely. By the time I was emailing Vanden for help, I was struggling to even clear one of the hint puzzles. When I messaged him, my stress was exacerbated by strange things I’d found in the game’s various hubs: suspicious passageways, odd blocks and some intriguing locked doors. These were hints of secrets, I realized, yet I sensed I needed to play further into the game to understand them. I was stuck, but I wanted to see the rest. Vanden offered to give me the exact solutions, to solve those three puzzles for me, but first he gave four nudges. Three were for the specific puzzles, but his best hint was for all and ultimately proved to be a terrific life lesson: go back a bit, he suggested. Do the puzzles right before those. They were teaching you something. Focus on what their lesson was. Don’t rush, in other words. Recognize what you’ve accomplished. Build off of that. I did so. I absorbed the lessons. I saw those three tough puzzles in a new light. I solved one the night after he wrote back, another the morning after. I left a third for another time. I didn’t need it and I wanted to plunge deeper into the game, closer to the moment when I was sure I could start unlocking those secret doors. My head stopped hurting, too. He Who Watches feels special. It’s got more challenges, and I’m undeterred. It’s also nice that its puzzles don’t have a surprise second health bar. Game File is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe Item 2: In brief… 🎮 Ilari Kuittinen, co-founder of Helsinki-based Housemarque (Super Stardust HD, Resogun, Returnal) has stepped down from his studio, after a 31-year run, saying he’s “ taking a rare and well-deserved breather.” Sony purchased Housemarque in 2021, and the studio released its newest game, Saros, earlier this year. Trend-watch: Sony-owned Insomniac saw its founder Ted Price retire in early 2025, after a 30-year run. Sony-owned Sucker Punch saw its co-founder Brian Fleming step down in late 2025, after a 28-year run. 👀 Savvy Games Group chief Brian Ward is stepping down from the Saudi-government-backed games conglomerate, Bloomberg reports. No word on a long-term successor yet, nor if this augurs any closer integration between Savvy and Electronic Arts. EA was purchased by a consortium led by Saudi’s sovereign wealth fund, in early August. 💿 Square Enix has denied a Japanese media report that the company was planning to go private. In very different Square Enix news, the company revealed that the physical version of Final Fantasy VII: Revelation, the April 2027 finale of its radical remakes of FFVII, will include just one disc in its case, requiring a download for the rest. The first two FFVII remakes in the trilogy shipped on two discs apiece. While Square hasn’t commented on the reason why, a feature here in Game File last month broke down the processes and costs for physical games. The expense of a second disc comes out of the game publisher’s cut of each copy. Square could be saving in excess of a dollar a copy with this move. 🟥 Grand Theft Auto IV’s 26-minute ’‘extended look” gameplay showcase that premiered on Netflix last week—and was exclusive to the platform for six hours—drew 31.1 million views in its first four days, the streamer announced. (Netflix calculates views as total viewed hours divided by runtime; excludes replays). Compared to YouTube: The version that later premiered on Rockstar Games’ YouTube channel is sitting at 22.6 million views after seven days (YouTube’s more liberal counting system tallies anyone who viewed even just the first few second.) Compared to other Netflix mega-events: Last year, Stranger Things season 5’s first four episodes, which debuted together, drew a record 59.6 million views in their first few days on Netflix. For more of a live event type of comparison, in 2024, the Jake Paul vs. Mike Tyson boxing match drew 65 million concurrent streams as it happened; Netflix didn’t release live numbers for the GTA premiere. ☁️ Microsoft is changing the way it charges for cloud gaming, from an all-you-can-play offer for Xbox Game Pass subscribers, to caps for the subscription’s tiers (15 hours per month at the Ultimate tier), with the option to buy more hours in bulk beyond that. The company also said it will sell batches of cloud gaming time to non-Game Pass subscribers, details TBA. Microsoft does not release figures for cloud game usage, but said its changes would impact 4% of Game Pass subscribers. That implies a small portion of its subscribers exceed the 15-hour monthly mark. Cloud gaming offers the promise of playing games without the need for console or PC hardware, as the graphics, sound and controls are all transmitted via a streaming connection. Microsoft notes that its cloud gaming service can be used for a range of things, including Xbox One players using it to play newer-gen Xbox Series games via streaming; and people playing games via streaming, sans any console, “in markets where consoles are less available or affordable.” (That latter use case could even be relevant in Xbox’s home market soon, as console prices keep rising.) Take Two Interactive chairman Strauss Zelnick recently told investors he expects game-streaming tech, which has been only lightly adopted by gamers for over a decade, to be widely used within three years. 🇺🇸 The Trump administration has published multiple agenda-driven web video games to the White House’s website. One called Build The Wall partially clones Tetris as it asks players to build a wall to stop a “zombie border siege.” While actual Tetris involves clearing lines and basically not building a wall, the Trump-backed game omits the line-clearing aspect (and omits the fun as well). On Instagram, Tetris’ official license holders reacted by stating: “At Tetris we believe in the power of connection and bringing people together, not dividing them. To our fans everywhere: we love you, we see you, and we’re grateful to have you in our community. P.S. The Tetris Company was not involved in the creation of ‘Build the Wall.’ P.P.S.: We take copyright infringement very seriously.” Item 3: The week ahead: Tuesday, September 8 Game releases: Mewgenics (consoles; already out on PC) is an engrossing tactical combat game (with cats). It is getting its console port right after the release of a Star Wars tactics game and ahead of the late September launch of a Fire Emblem–a bumper crop for the genre. Wanderburg (PC) is an auto-shooting game like Vampire Survivors, but you’re a castle on wheels Event: Nintendo airs a 30-minute Nintendo Direct (10am ET) focused on The Legend of Zelda’s 40th Anniversary. A full unveiling of the upcoming remake of The Legend of Zelda Ocarina of Time is expected. Will they also show footage from next year’s live action movie? Tease the next big 3D Zelda? Add Link to Mario Kart World? Shadow-drop a new Tingle game? Wednesday, September 9 Event: Nintendo again! On this day, they’re planning to air a 45-minute Nintendo Direct (10am ET) focused on games heading to Switch 2 this winter. Followed by about an hour of gameplay demoed by the company’s Treehouse team. The September Nintendo Directs are always some of the year’s biggest. The 2025 edition included announcements of several major 2026 Switch/Switch 2 games: Pokémon Pokopia, Mario Tennis Fever, Yoshi and the Mysterious Book, Fire Emblem Fortune’s Weave. Plus the return of the Virtual Boy. Friday, September 11 Game release: EA Sports NHL 27 (PlayStation, Xbox) is a hockey game. There are so few that an indie studio that had been making wacky sci-fi first-person shooters just announced that, hey, we’re going to fill a void and make a hockey game for PC.
If Not AOC, Then Who?
What do we do if Alexandria Ocasio-Cortez decides not to run for president in 2028? It makes sense that many leftists are currently focused on calling on her to take the plunge. But the absence of a Plan B, or even much discussion about one, is worrisome. The US Left can’t afford to sit out the 2028 Democratic primary. Doing so in 2024 was a disaster: our abstention self-marginalized anti-billionaire politics from the national conversation and thereby gave space for Harris to pivot to the “center” (i.e. the policy preferences of corporate donors, establishment hacks, and AIPAC). Abstaining again in 2028 would be especially counterproductive since democratic socialists now have so much wind in our sails after our big electoral wins in NYC and across the country. It’s time to deepen our momentum, organization, and profile — not to retreat to the margins during America’s single most important political battle. Even more urgently, keeping Trump’s heir out of the White House may depend on whether leftists and the party’s unprecedentedly disenchanted base can force the Democrats in 2028 to take meaningful steps towards a very angry electorate and away from the corporate-consultant blob that has helped get us into this mess. If AOC (or Abdul El-Sayed) decides to run, we should enthusiastically go all in — DSA, fighting unions, progressive organizations, anybody who wants to see a better country and world. But, as far as I can tell from the outside, there’s a real chance that she doesn’t run. And we can’t just wait and see where she lands before discussing alternatives, since AOC’s decision might not come for quite a while, giving us little time to scramble at the last minute for an alternative. So we should start working now on getting a Plan B into motion, and recruiting a good candidate for the role. Here’s one idea. A Worker Candidate — Healthcare Not Warfare 2028 Run a union worker Leftists don’t need to always run experienced, well-known socialist politicians for high-profile contests. Lula, the current president of Brazil, began his political career as a metalworker militant who ran for governor of São Paulo in 1982. Especially in an era of institutional distrust, it’s possible to lean hard into making a virtue out of a working-class comrade’s distance from traditional politics. Experience abroad shows this clearly. Rachel Kéké — an immigrant hotel housekeeper — ran as the candidate of the leftist La France Insoumise for the National Assembly in 2022, defeating a former minister in the Macron government. The Belgian Workers’ Party in 2024 won its first Flemish seat in the European Parliament by running Rudi Kennes, a three-decade veteran of the Opel Antwerp auto factory. And the South Korean left in 2024 elected auto factory militant Yoon Jong-oh to the National Assembly. If a tiny French Trotskyist organization could make such a big splash in 2002 and 2007 by running a well-spoken postal worker cadre for office — the LCR’s Olivier Besancenot received about 1.5 million first-round votes — it’s not far-fetched to expect that a much larger organization like DSA, together with a broader coalition of the willing, could run this playbook to much greater effect. This is especially the case now that almost everything DSA does immediately becomes a national news story. We don’t need someone famous. We just need a charismatic, well-vetted comrade who works in a relatable working-class occupation — e.g. a nurse, or a worker in auto or construction, or at Amazon. Ideally, they have a compelling personal story or have been part of a significant strike or organizing campaign, but even those criteria don’t necessarily have to be deal-breakers. The main lesson of the Graham Platner fiasco isn’t that we should never run outsiders for office (the momentum he generated proves the exact opposite), but that they need proper vetting and roots in democratic membership organizations like DSA or a union. In any case, there are plenty of ways to signal working-class authenticity without resorting to Platner’s or Fetterman’s performative macho white dude affect. Better to have someone of any gender or race who is actually working class, with firm socialist politics, who can speak from experience about the day-to-day realities of trying to make it in America. There are surely at least a dozen such worker leaders in our movement nationwide; we just need to sit down and convince one that their personal sacrifice for the greater good (i.e. running for president) will be worth it. If we don’t have a frontrunner with a clear path to victory like AOC (or El-Sayed, were he to win Michigan), it’s better to embrace someone whose decision to run in itself tells an important story about US politics: workers have been iced out of our political system and are now fighting hard to break back in. The novelty factor, in itself, would generate press. And it would help democratic socialists signal to the American public that our movement’s defining characteristic is that we represent workers vs. bosses, not that we’re “the furthest left” force in American politics — a framework that suggests we’re extremists and that helps our opponents isolate us from non-college-educated voters by highlighting extremely minoritarian (and misguided) stances included in DSA’s platform like prison abolition. We don’t get many opportunities to speak to tens of millions of Americans; we need to use these opportunities to broaden our appeal beyond those already in our orbit. Run to polarize the 2028 primary and general election around “Medicare for All” and “End All Aid to Israel” Why should democratic socialists use our scarce time and volunteer energy to run in a race we’re unlikely to win? Running in the 2028 primary will not only recruit us scores of members (imagine having a DSA member intervening in eleven nationally televised primary debates!), but — more importantly — it has the potential to be a decisive leap forward towards an America that finally starts meeting the needs of working people instead of HMO and AIPAC billionaire donors. Polarizing litmus tests are good, actually — as long as they’re fought for around popular demands in a thoughtful and strategic way. Democratic socialists have an opportunity and responsibility to make the 2028 primary about two of the most widely and deeply felt issues of our moment: domestically, America’s unaffordable, family-breaking private healthcare system; internationally, America’s bankrolling of the genocidal Israeli government. A candidacy relentlessly demanding Medicare for All and No Aid to Israel would pose a sharp question to the American public, the Democratic Party, and other 2028 contenders: Why can’t working-class Americans afford their healthcare bills at the same time that our government is spending at least $3.8 billion yearly to fund an Israeli regime eager to keep dragging us and the world into more forever wars? Our candidate can agitate around a clear message: I’m running to give voice to all those working-class Americans who are struggling to get by. No matter what state you live in, who you voted for in 2024, or your background, it’s just plain common sense that we should make healthcare more affordable instead of spending billions to enable the Israeli government’s crimes against humanity. Every American who agrees with me when I say ‘Medicare for All — Not a Dime for Israel’ should make their voice heard through our campaign. Racking up a high number of votes for this agenda is a crucial way to show the Democratic Party and the country as a whole that we can’t keep supporting a life-destroying status quo in our healthcare system and in the Middle East. If we want to defeat Trump’s handpicked crony this coming November, we need to speak to ordinary people’s frustrations with the status quo — we can’t keep running the same tired establishment playbook and expecting a different result. For those highly engaged voters (or skeptical media pundits) who raise concerns about wasting votes, the candidate can take the time to explain why this isn’t the case: It's true I'm a long shot. But Zohran showed anything can happen — and voting for me is a pragmatic decision, not a wasted vote. In a heavily divided primary with a large number of candidates and a very discredited establishment, my campaign and the democratic socialist movement I’m part of are positioned to be the balance of power if we can get a large enough number of voters to loudly demand Healthcare Not Warfare. Since there’s a high chance no candidate will get to 50% in the 2028 primaries, the Democrats could very well be heading towards a brokered convention, where a disciplined, principled current committed to breaking from corporate donors can play an outsized role in making sure whoever ends up with the nomination raises the banner of single-payer healthcare and cutting aid to Israel. How exactly could running this sort of primary campaign be translated into actual leverage? The candidate replies: I’m aiming to win the nomination, but if I don’t achieve that goal, my delegates and I will pledge our support in the primary only for a candidate who agrees in the general election to call for ending US aid to Israel and establishing Medicare for All at home. These are both winning planks for November against the Republicans, and the only reason I can see that other candidates haven’t yet raised these is that they’re still more scared of the donor class than they are of the people. That needs to change. If Democratic leaders want young people, workers, and those sick of the status quo to show up in big numbers this November, they need to meet us halfway. In this way, whether or not we win the nomination, democratic socialists could still play a decisive role in 2028. The approach outlined here is similar to what radical left parties periodically do in parliamentary democracies like Denmark or Sweden when they find themselves as a junior partner electorally: they dictate to center-left allies what their terms would be for enabling, from the outside, the formation of a new government. Running a Healthcare Not Warfare candidate will make it significantly more likely that the 2028 Democratic candidate stands with ordinary people instead of donors on these two pivotal planks — a shift that will, in turn, make it easier to achieve the existentially important task of defeating MAGA’s heir in November. Look at how much popular traction the Uncommitted movement got with a similar strategy in 2024 with so little on-the-ground organization — surely we can go much further with a working-class candidate, a much bigger DSA, and a Democratic Party base that has turned against Israel, billionaires, and the hacks in the party who prop them up. Objections and Responses There are quite a few objections to this proposal, many of which are very reasonable. Let me briefly address each here. Socialists shouldn’t waste their scarce time and resources on a candidate unlikely to win. This is a good rule in general, but there are important exceptions, and presidential elections are the main one. There’s no other political contest that gets anything like this level of attention and engagement in the US, so if you want to be a serious force in nationwide politics, you basically can’t afford not to have a candidate. Far from being an overall drain on our resources, running a dynamic 2028 campaign would likely result in a major net increase in members and capacity. In that spirit, the very pragmatic Milwaukee sewer socialists — who mostly ran to win at home in Wisconsin — supported Eugene Debs’s propagandistic presidential campaigns, which they correctly saw as a crucial means of recruiting to the Socialist Party. And now that the (post-1972) Democratic Party gives more weight to voters than party operatives in determining its presidential nomination, it’s become feasible to run as a socialist in primaries without risking a split in the vote against the Republicans in November. We should just support Ro Khanna, who also supports M4A and ending aid to Israel. I like Ro Khanna and would enthusiastically vote and canvass for him for president were he to get the nomination. But there’s a real chance that Khanna’s bid for office doesn’t catch on; he lacks the grassroots organizational apparatus of DSA, and, because he’s not really of the Left or labor (despite his current very good positions), he’s unlikely to generate the type of buy-in and enthusiasm that a DSA candidate would. Moreover, Khanna would be running as an individual, not as someone emerging from and building an organized movement — his campaign would not likely focus on building sustained power from below, nor would it likely polarize the 2028 primary around M4A and Israel aid. Also, and not inconsequentially, the poor electoral results for Tom Steyer and Saikat Chakrabarti this spring suggest that there’s perhaps not much of a popular appetite in the Democrats’ base to support multi-millionaire class traitors. Much better, then, to run a democratic socialist worker. If by Super Tuesday it’s clear that our campaign has captured the main left lane in the primary, Khanna should drop out and endorse us; conversely, if he’s managed to capture this lane, DSA could at that point transform its Healthcare Not Warfare campaign into an independent campaign to elect Ro Khanna focused on winning (and holding him accountable to) those two critical planks. An openness to consolidating an anti-corporate candidate by Super Tuesday is a necessary component of the strategy I’m proposing, because otherwise we risk splitting the primary votes (per DNC rules, you need at least 15% to get any delegates to convention). We should convince Shawn Fain or Sara Nelson to run for president instead. I agree it’d be great if either ran for president, but my impression is that there’s basically no chance that will happen in 2028, because both are firmly focused on their union leadership duties (and, in the case of Fain, beating back a federal witch-hunt). This plan could have DSA end up endorsing a non-socialist, non-DSAer for president. That’s against our principles. It’s a myth, and bad politics, that socialists can’t ever support candidates who aren’t members of their organization and/or who don’t call themselves socialists. Bernie, after all, was not a DSAer. And socialists abroad regularly make tactical decisions about if and when supporting a non-socialist candidate makes sense from the standpoint of promoting working-class interests, defending democracy, and building a socialist current in the process. If your “principles” prevent you and DSA from helping achieve the momentous stride forward of having a Democratic candidate in 2028 run on Medicare for All and ending aid to Israel, those “principles” are not very good. In fact, they actively undermine the causes they’re aiming to support. We shouldn’t confuse principles (working-class independence) with tactics (the best way to advance this in a given context). In my recent article in Catalyst on Milwaukee sewer socialism (I’ve uploaded the full text here), I have a section showing how similarly rigid rules about endorsements had to be eventually discarded in the old Socialist Party nationwide after WWI because they tended to lead socialists towards a sectarian abstention from broader anti-corporate electoral insurgencies beyond their party’s control. For what it’s worth, Lenin’s current under Tsarism frequently backed bourgeois candidates in the second round of elections and Lenin himself advocated that British Communists support Labour Party candidates; for his part, Kautsky in 1912 defended the German Social Democracy’s electoral pact with the Progressives in the runoffs. If our predecessors could be that tactically flexible in an era before the limits of their underlying strategic assumptions about intransigent class politics in capitalist democracies were made clear, then there’s no need to insist on being more tactically rigid today, after a century’s worth of accumulated experience. As long as you maintain your own political profile and organization while engaging in broader coalitional efforts, there’s zero inherent contradiction with the principle of political independence. More specifically: DSA and allied organizations should enthusiastically back Michigan’s Abdul El-Sayed if AOC declines to run and he throws his hat in the presidential ring. (I have no idea whether he’s even considering this path, but there’ll surely be chatter about it if he wins his Senate race this November.) Though he doesn’t call himself a socialist, El-Sayed is a principled Berniecrat, a consistent fighter for Palestine, and would have a clear path to victory. Using the “s” word is far less important than whether a campaign fights for the interests of working people at home and abroad — and whether actively participating in it can grow the organized working class and the Left. And the same holds true if a less anti-corporate candidate were, under pressure, to adopt our two Healthcare Not Warfare planks. Why focus only on M4A and ending US aid to Israel? There are so many other important issues to fight for! Obviously there are many pivotal issues beyond healthcare and Israel, and our candidate — and especially our organizations involved in the campaign — should be prepared to speak on all of these. But one of the crucial lessons of the Zohran campaign is that you have to relentlessly stay on message if you want to cut through the noise. CNN and Fox News will want to talk about DSA’s views on Venezuela and abolishing the police; we need to polarize the debate on the widely and deeply felt issues that are most favorable for growing the Left, isolating our opponents, and winning urgently needed changes. One recent poll found 63% of Americans support M4A, even after respondents were told it would raise taxes and eliminate most private insurance. Another found that 74% of Democrats opposed “providing additional economic and military support to Israel.” And focusing on two demands sets the campaign up to be able to leverage its mandate to make clear demands upon other candidates (which we can’t do if we run with a laundry list of demands). Talking about a Plan B is a distraction from the more urgent work of building a groundswell of calls on AOC to run. It’s good for DSAers, fightback unionists, and progressive organizations to try to convince AOC to run. This proposal is meant as an addendum to pushes to get AOC to run and to line DSA up for that eventuality, not as a substitute for those efforts. But I’m worried that, like in 2024, we might eventually find ourselves with no time to find a viable alternative unless we start now. DSA, with its abundant deliberative processes, moves slowly. The approach you’re suggesting seems risky. If a DSA candidate runs but doesn’t receive a lot of votes, it’ll undercut our current national momentum. This is a reasonable worry. All big initiatives always carry risks, and the path I’m advocating is certainly not free of these. That said, if a Healthcare Not Warfare campaign doesn’t catch on like we envisioned, it’s a pretty simple pivot in March 2028 to explain that in the future we really need our standard-bearers like AOC, or our growing bench of national electeds, to step up in presidential primaries. Moreover, it’s worth noting that the risk of my Plan B proposal is on the whole significantly less than the risk entailed by an AOC run. We should all back her if she runs, since she’d fire up the base, provide huge openings for organizing, have a real shot at defeating a Republican (especially given MAGA’s unpopularity), be a great standard-bearer for our issues, particularly M4A and ending aid to Israel, and be less willing than mainstream candidates to capitulate to any anti-democratic machinations from Trump and co. But it’s also true that AOC could lose, which would be a devastating setback to the country and the Left. Or she could win and find herself unable to pass much, if any, of her (and our) program, a demoralizing prospect. So if we’re willing to take these risks with AOC (and we should), it follows that we should be able to take lower-level risks if she doesn’t throw her hat in the ring. What’s Your Plan B? The US Left is facing its biggest opening for growth and influence since the 1930s. Initiatives that would have been impossible a few months ago are now potentially feasible. A President AOC is now easier to imagine, but if she ultimately assesses that the risks aren’t worth the potential reward in 2028, we can’t be left flat-footed. So even if you’re not convinced by the specifics of my Plan B, I hope you can at least agree that it’d be a disaster for us not to have one — and not to start getting a Plan B into motion. Hopefully other organizers and political currents will put forward their responses and alternative ideas ASAP. The stakes are too high to just wait and see. More As always, please share this article widely. This newsletter depends on all you lovely people to get the word out. Thank you! I uploaded the full text of my Catalyst article on Milwaukee sewer socialism here (it’s updated and 3x as long as the version published earlier on this substack). If you haven’t yet, please make sure to subscribe to Catalyst here; I read every issue front to back, highly recommend. 🌹🎵 Ivan Lins (arranged by Arthur Verocai) — “General da Banda” 🌹🎵
The Streaming Bubble Keeps Popping
(Welcome to the Entertainment Strategy Guy, a newsletter on the entertainment industry and business strategy. I write a weekly Streaming Ratings Report and a bi-weekly strategy column, along with occasional deep dives into other topics, like today’s article. Please subscribe.) One of my rules is “never bet against PE [private equity]”. This isn’t to say that I agree with their methods or tactics or strategies or financial innovations (aka shenanigans). Just that in the long arc of financial history, they’ve tended to return very good returns to their shareholders. That said, their forays into entertainment have left me...skeptical. I’ve written before about PE’s production company/celebrity production company roll ups, and those haven’t really worked out. Some PE groups tried to “roll up” an industry that isn’t “roll-up-able” for lack of a better term—since the barrier to entry for starting a production company is very low—often paying top dollar for celebrity production companies at the height of streaming valuations. So yeah maybe we can bet against PE when they venture into filmed entertainment. I’ve also noted that studio lots seemed to be a bubble back in 2022. Sure enough… Yikes! That sure sounds like some private equity guys lost some money. The only caveat being that holding a lot of land in Los Angeles is always valuable. But that wasn’t the pitch and likely real estate investment companies overpaid for said land. I mean, Goldman Sachs had to repossess it after all. (Technically, I’m not sure these buyers were traditional private equity, but PE was buying LA studio lots earlier this decade.) All to say: I call out bubbles when I see them, and they often pop. Anywho, my most famous bubble call was the “streaming bubble”, specifically that the major streamers were making too many shows and films. And that it could all come crashing down. Notably, I made this call before the strikes of 2023, but those strikes potentially accelerated the popping. And now it has indeed popped. One of my most popular articles of last year made that exact case. A year later, it’s time to update that analysis. The bottom line is that the streaming contraction continues. Credit where credit is due, Luminate—an analytics company whose data I use weekly in my streaming ratings report—called this out in a recent report I saw highlighted in both Bloomberg and The Hollywood Reporter. But I’m adding my data to their look, including streaming films and kids TV shows. I’ll also provide some other data that all tells the same picture. Let’s dive in! Subscribe The Number of Streaming TV Shows, Films and Specials Is Declining AGAIN in 2026 Let’s get right to the data, starting with my dataset. Specifically, each week I track every “notable” title to come out on streaming. I try to grab every title on every major streamer, meaning if it’s a first run or original or exclusive, I track it. This helps me call out the “dogs not barking” and the misses/flops each week. So that will be our first series of charts, my weekly collection data. Let’s start with the raw TV show, film and special data, by week. Note: As the years have gone by, I have actually added more sources (going from just one source to five sources) to find new titles, meaning that this data collection is, likely, more accurate now than in the past. Which means my team and I are more likely to have undercounted titles in the past rather than today. I like to cut the data several different ways. Here’s that same look, by quarter, to more sharply show the trends: And here’s the data just looking at the first half of each year. That’s the same data, just cut three different ways with three different time periods. Note that films, TV shows, and specials are all down by 25%, 42% and 17%, respectively, from 2022. The good news is that the number of films and specials actually increased year-over-year from 2025 to 2026, but not enough to offset the 25% decline in TV shows. Now, I also look at kids shows, and here they are pulled out on their own: And here’s that by quarter: This really seems to be a notable area of pullback by the streamers. Kids shows are down 80% from 2022. This reflects some formerly Disney+ shows going to cable TV first, but also pullbacks from HBO Max, Prime Video and Netflix. I also categorize the TV shows by English language and non-English language, and that also reveals a lot of where the drop comes from: And here is by the half-year: In this look, the good news is the drop for English language shows seems to have stabilized in the last year. But the bad news is the long term decline from the 2022 peak. Other Data So I wasn’t the only person to notice this decline in the number of shows. I have a few other data sources for this, though most only go through the end of 2025. As I said, Luminate beat me to the punch, so let’s look at their data. We’re just getting started with this issue, but the rest is for paid subscribers of the Entertainment Strategy Guy, so if you’d like to find out… What TV loss streaming didn’t make up for... How Netflix has changed their programming slate over time… Ten more images charting the decline of production... What type of films aren’t being made anymore... The good news (for movie theaters)… And more... ...please subscribe! We can only keep doing this great work with your support.
Shaky political "science": breakdown of the peer review process
[This article is the latest in our “Shaky Political Science” series. In our 38 page report, “Shaky political ‘science’ misses mark on ranked choice voting,” released in December 2025, we examined over 40 studies focused on ranked choice voting (RCV). Many of these studies suffered from puzzling research methodologies, poorly constructed surveys, unrealistic simulated elections, cherry picked data and faulty analyses that often were contradicted by results from real-world elections (a link to our full paper is here). Here are links to the first four articles in our series: the first one summarizing the overall results of our report; the second one analyzed a study co-authored by University of Minnesota academic Larry Jacobs, which was one of the most error-prone of all the research we reviewed in our report; a third article showing an admirable example of two political scientists holding two other political scientists accountable for their flawed, sloppy research; and a fourth article that analyzed seven studies sponsored by New America’s political reform program that deployed flawed methodology based on “simulated” elections instead of real-world elections, which when combined with some basic misunderstandings of RCV, produced odd and misleading results]. In this fifth article in our “Shaky Political Science” series, we examine two flawed studies on the impact of RCV on voter turnout authored by San Francisco State University professor Jason McDaniel. McDaniel’s research in these two papers stands out for its puzzling methodology, inadequate and cherry-picked data, and faulty analyses that not only misunderstands ranked choice voting but reflects a basic misunderstanding of politics in general. The first of the McDaniel studies actually passed through a peer review process and then was subsequently published by the Journal of Urban Affairs, despite its obvious errors. It is perplexing that a jury of peer political scientists did not catch the errors and send it back to McDaniel for a rewrite or outright rejection. This calls into question the credibility of the peer review process itself. The peer review process is supposed to represent the intellectual pinnacle of the profession – a way for academics to certify the quality of each other’s work, and to act as a check against sloppy research. In the case of the McDaniel study, the peer review process failed. Yet his flawed study has been cited dozens of times by other political scientists in their own flawed works — an unvirtuous circle of academic research. Here is the first Jason McDaniel study that we analyzed: “Writing the Rules to Rank the Candidates: Examining the Impact of Instant-Runoff Voting on Racial Group Turnout in San Francisco Mayoral Elections,” by Jason McDaniel, 2016, Journal of Urban Affairs. This study covered San Francisco’s use of ranked choice voting (RCV) in mayoral elections, purporting to show that RCV reduces voter turnout overall, increases voter errors, and has a negative impact on “marginal populations,” i.e. racial minorities, mainly as the result of the alleged complexity of RCV asking voters to rank candidates and having to be familiar with multiple candidates and related dynamics. McDaniel reached these conclusions by examining five San Francisco mayoral elections taking place in odd years (i.e. where there are no federal or state races on the ballot) from 1995 to 2011, comparing three earlier non-RCV races (in 1995, 1999 and 2003) with the later two RCV races (in 2007 and 2011). Only two RCV races were studied, and as McDaniel acknowledged, only the contest in 2011 was “competitive” (that contest in 2007 was the first mayoral election using RCV, and incumbent mayor Gavin Newsom won easily in the first round of counting, earning 74% of first choices and far ahead of the second place candidate who had just 6.3% of first choices). Yet even the allegedly competitive contest in 2011was won by an incumbent with a landslide margin of nearly 20 points. No other elections were studied by McDaniel, even though San Francisco elects 17 other offices besides mayor with RCV, including the 11 seats on the Board of Supervisors (the name for the city council) in even years, and six other citywide offices in odd years. Previous academic research has found that competitive elections can stimulate greater voter involvement and higher turnout. For unexplained reasons, this study ignored at least 34 other RCV races between the years 2004 to 2011, many of which were very competitive and had high voter turnout. Instead, McDaniel’s analysis based its conclusions on data from only a single, marginally-competitive mayor’s race using RCV. These 2007 and 2011 mayoral elections were the only two RCV elections that McDaniel used as the sole basis for his small data set to determine the impact of RCV on voter turnout. Despite having no meaningful data to inform his study, McDaniel nonetheless concluded that RCV lowered turnout when compared to the previous three mayoral non-RCV elections using two-round runoffs. Then, using the same inadequate database of San Francisco elections, McDaniel also found that “Instant-runoff voting is associated with a significant decline in Black voter turnout and White voter turnout…On average, Black voter turnout decreases by 18 points and White voter turnout decreases by 16 points …” But how could he credibly arrive at that conclusion when his study was based on voter turnout in two RCV contests that were not even remotely competitive? Beyond that, McDaniel tacitly admits that the key determining factor driving Black voter turnout was not RCV vs non-RCV elections, but the presence of a viable Black candidate. McDaniel wrote, “Black voter turnout was about 17 percentage points higher for the two elections that featured an African American candidate compared to the three elections that did not” — but one of those three elections did not use RCV. The same dynamic was present for White voter turnout, in which higher turnout was not correlated with RCV vs non-RCV elections, but the presence of a viable White candidate. Beyond that, McDaniel’s voter turnout results, whether correlated with the race of the candidates or RCV vs non-RCV elections, showed no clear patterns. In his Figure 1, he shows low Asian voter turnout in the non-RCV elections in 1999 and 2003 when there was no viable Asian candidate running, but in the 2007 mayoral election using RCV, there is still no viable Asian candidate, yet Asian turnout increases dramatically. For Latinos, voter turnout is virtually the same in 1999 (no viable Latino candidate, non-RCV elections), 2003 (leading Latino candidate, non-RCV elections), 2007 (no viable Latino candidate, RCV elections) and 2011 (leading Latino candidate, RCV elections). Meanwhile White voter turnout increased dramatically from a low in the 2007 mayoral election, in which the leading candidate was a white incumbent, to the 2011 election in which there were no leading white candidates that finished any higher than sixth place with 5.6% of the vote. Clearly other factors were determining voter turnout than either the race of the candidates or RCV vs non-RCV elections. Despite these inconsistent findings, McDaniel boldly marches forth to render conclusions about the impact of RCV on different groups of racial voters, even though his conclusions are not remotely supported by his data. This looks like the worst kind of “data cherry picking” to cover up a puzzling conclusion based on a poor data set involving only two RCV elections, neither of which were remotely competitive. McDaniel’s cherry picking didn’t stop there. In calculating voter turnout for the three non-RCV mayoral elections in 1995, 1999 and 2003, he failed to mention that in each of these years there was a two-round runoff election, with a first round election in November in which no candidate garnered a winning majority of the vote, followed by the decisive runoff in December. So there were two non-RCV elections in each election year. But McDaniel only decided to include one of the two elections in his data set, apparently the December runoff in each year (which had higher turnout in two out of the three mayor’s races, and lower turnout in the third). So he ignored the November 1995 election, the November 1999 election when turnout was only 45 percent, and the November 2003 election when turnout was only 45.7 percent. This is flawed methodology based on blatant cherry picking, since McDaniel simply omitted three elections worth of data that did not fit within his frame. His study should never have made it through a peer review process, and it is an embarrassment to political science that not only was it approved for publication, but it has been uncritically cited by dozens of other political scientists. Share A “study” without wider context or credible data Lacking a data set that included even a single competitive RCV mayoral race, McDaniel declared that RCV lowered voter turnout in 2011, when turnout was 42.5 percent. But he also ignored -- or did not realize -- that voter turnout in mayoral elections during 2011 had sharply declined in all big cities, including far more steeply in nearby Los Angeles and other cities. In fact, San Francisco’s election in 2011 had the second highest turnout of any mayoral election in the nation’s 22 largest cities from 2008-2011. Was 2011 just a particularly low turnout year all across the country, perhaps because the country was still reeling from the aftermath of the housing market collapse in 2008-2010? McDaniel ignored this factor and did not search for alternative explanations beyond his flawed thesis. Furthermore, while McDaniel suggested that RCV leads to more “overvotes” – which is an invalidated ballot in which a voter, in either a non-RCV or RCV race, makes a mistake and selects two candidates for a single choice – he failed to address that in the June 2012 US Senate primary contest, which used the typical “one-choice” plurality method and did not use RCV, there was more than five times as many overvotes in San Francisco and Oakland than in the RCV mayoral elections in those cities. Additionally, turnout had been significantly higher and more representative in RCV elections for the San Francisco Board of Supervisors races -- that McDaniel didn’t even bother to study -- as compared to previous “delayed two-round runoffs” for those offices, often higher by as much as 40% in the RCV races. Perhaps the most surprising irony was that an earlier study co-authored by McDaniel (Overvoting and the Equality of Voice under Instant-Runoff Voting in San Francisco by Francis Neely and Jason McDaniel, 2015, California Journal of Politics and Policy) found that overvotes are often more common in lower-income precincts, but that the pattern of overvoting is similar in both RCV and non-RCV contests. Wrote Neely and McDaniel: “[Overvote] errors appear to be a function of complexity in general and not IRV per se…. the pattern of overvoting is similar in both IRV and non-IRV contests…Even among votes cast in a top-of-the-ticket race like U.S. Senate, overvoting occurs in an uneven fashion” (IRV, or ‘instant runoff voting,’ is another name for ranked choice voting in single-winner contests). So McDaniel not only contradicted his own conclusion in his earlier study with Neely, but the earlier study had been an important corrective for one of the most frequently made mistakes made by anti-RCV political scientists – a continual failure to provide crucial context by pairing a comparative analysis of real world RCV elections with an analysis of real world plurality or other electoral methods. Unfortunately McDaniel’s earlier study does not get cited much by other researchers, and McDaniel did not even cite it in his own flawed study analyzed above. Incredibly, McDaniel cherrypicked McDaniel, even as his flawed study has been cited dozens of times. Despite being a textbook example of poor political “science,” featuring both inadequate as well as cherrypicked data, the shelf life of McDaniel’s study has been extended endlessly by other political scientists and media outlets who have cited it uncritically. Equally troubling, the Journal of Urban Affairs put this paper through its peer review process and published it despite its numerous flaws. McDaniel gets it wrong, Part II Not content with producing one shoddy, confused paper, McDaniel produced another one a few years later, here is the link: Electoral Rules and Voter Turnout in Mayoral Elections: An Analysis of Ranked-Choice Voting by Jason McDaniel. San Francisco State Univ. (McDaniel, 2019). In this paper, McDaniel extended his previous flawed research focused on San Francisco, about voter turnout and voter impact, to multiple RCV cities through the 2018 elections. In his data set he included San Francisco, Minneapolis, St Paul, Oakland, Berkeley, San Leandro and Santa Fe. He then compared their turnout in mayoral elections to non-RCV cities, and to mayoral elections in those cities held before first RCV usage, and concluded “a significant decrease in voter turnout of approximately 3–5 percentage points in RCV cities after the implementation of RCV.” But he replicated many of the same mistakes as his first study presented above. While in this study McDaniel attempted to distinguish between various factors, such as odd year vs. even year elections, competitive races vs. noncompetitive, contests with open seats and other factors, the analysis is superficial and inadequate. He attempts to incorporate so many often-conflicting variables that his conclusions are once again unsupported by the actual data, and ultimately unconvincing (this second study never underwent a peer review process). For example, McDaniel declared that for RCV cities “in elections that are more competitive than average (margin of victory less than 20%), there is no significant RCV effect on voter turnout.” Yet McDaniel is a professor in a city, San Francisco, in which the 2018 special election using RCV to fill a mayoral vacancy was exceedingly close, decided by a margin of only one percentage point, and resulted in the second highest turnout in San Francisco history for a mayoral election. Illustrating the complexity involved in assessing the critical factors in turnout, that special election was held on the same election day as federal and state offices, which tends to boost turnout. Yet more voters participated in that RCV mayoral election than voted on the same day in the non-RCV primary for governor in which former San Francisco mayor Gavin Newsom was running. Even more puzzling, in this study McDaniel inaccurately cited his own previous study from 2015 with Neely. McDaniel wrote, “Neely and McDaniel (2015) examine individual ballot image files within a natural experiment, and find higher rates of ballot errors in RCV elections compared to non-RCV elections” (emphasis added). However, we quoted from that 2015 study several paragraphs above, and in actual fact Neely and McDaniel actually concluded the opposite -- that the pattern of overvoting – which is the most common form of ballot error – is similar in both RCV and non-RCV contests. McDaniel also ignored the fact that in Bay Area RCV elections held in even years – at the same time as federal and state elections, when turnout is highest -- the “undervote” (that is, the dropoff in ballots cast for governor or mayor) has declined, which contradicts his thesis that the act of ranking ballots in RCV elections is too complex and presents a barrier for certain demographics of voters. To his credit, McDaniel actually acknowledged another contradiction in his own paper, which found a negative RCV turnout effect in some odd-year elections but not others, admitting that “may present a challenge to the theory that increased complexity [from RCV] will have a marginal negative effect on voter turnout.” But he tried to wallpaper over that contradiction by blaming it on an unlikely source -- bad election administration. McDaniels wrote, “This would suggest that the negative impact of RCV on voter participation may be alleviated by high quality election administration,” without presenting any data or even anecdotes to support such an unfounded theory. While competent election administration of any election, whether RCV or non-RCV, is crucially important, there is no credible research that supports the far-fetched notion that the quality of election administration has much of an impact on voter turnout in US elections, whether using RCV or otherwise. Upgrade to a $5 subscription The unvirtuous circle of academic research It is puzzling how McDaniel’s first study could have survived the peer review process for the Journal of Urban Affairs. The mistakes in McDaniel’s first study above should have been obvious to anyone with a basic understanding of ranked choice voting as well as elections in general, and elections in San Francisco specifically. It undermines faith in the credibility of the peer review process to think that a jury of peer political scientists did not catch the egregious errors and reject McDaniel’s sloppy research and paper. And it remains a black mark on political science that this thoroughly discredited study has been cited dozens of times by other researchers. Indeed, in our research paper it became apparent that many political scientists cite each other’s flawed research to pad their list of citations and to legitimize their own flawed research, creating an unvirtuous circle of political science. More broadly, McDaniel’s voter turnout studies, like other studies evaluating RCV’s impact on turnout, lacked a recognition that no credible opinion or research can claim that RCV increases voter turnout merely because so many voters are excited to vote for their favorite candidates instead of “lesser evil” candidates without worrying about spoilers. Instead, with respect to turnout, the advantage of RCV has always been that, in most places where it has been used, it has eliminated a pre-November primary or post-November runoff, which usually saw about half the turnout of the November general election. This is especially true during even years when there are state and federal elections on the ballot which substantially drive turnout more than local contests. Any study that does not adequately account for this non-November primary/runoff vs November general election dynamic, or odd year vs even year election factors, is simply not credible. Both of McDaniel’s studies largely failed on those grounds. Steven Hill and Paul Haughey Share Upgrade to a $5 subscription Thanks for reading DemocracySOS…your digital portal for the pro-democracy movement. Subscribe for only $5 per month to receive full benefits and to support our work. Subscribe
The Quintessential Trump-era Politician
In 2020, I wrote an article about collaboration and collaborators. The title was History Will Judge the Complicit, and it started with the story of two famous East Germans, Markus Wolf and Wolfgang Leonhard, who both landed in East Berlin just after the end of the war. Although they came from very similar backgrounds, the former stuck with the regime, eventually becoming its spy boss; the latter, horrified by the brutal reality of Soviet-style communism, escaped. By way of juxtaposition, I also wrote about Lyndsey Graham and Mitt Romney. These two men also made very different decisions when confronted by Donald Trump’s assault on the ideals that they had both grown up believing: A glance at their biographies would not have led many to predict what happened next. On paper, Graham would have seemed, in 2016, like the man with deeper ties to the military, to the rule of law, and to an old-fashioned idea of American patriotism and American responsibility in the world. Romney, by contrast, with his shifts between the center and the right, with his multiple careers in business and politics, would have seemed less deeply attached to those same old-fashioned patriotic ideals. Most of us register soldiers as loyal patriots, and management consultants as self-interested. We assume people from small towns in South Carolina are more likely to resist political pressure than people who have lived in many places. Intuitively, we think that loyalty to a particular place implies loyalty to a set of values. But in this case the clichés were wrong. It was Graham who made excuses for Trump’s abuse of power. It was Graham—a JAG Corps lawyer—who downplayed the evidence that the president had attempted to manipulate foreign courts and blackmail a foreign leader into launching a phony investigation into a political rival. It was Graham who abandoned his own stated support for bipartisanship and instead pushed for a hyperpartisan Senate Judiciary Committee investigation into former Vice President Joe Biden’s son. It was Graham who played golf with Trump, who made excuses for him on television, who supported the president even as he slowly destroyed the American alliances—with Europeans, with the Kurds—that Graham had defended all his life. By contrast, it was Romney who, in February, became the only Republican senator to break ranks with his colleagues, voting to impeach the president. “Corrupting an election to keep oneself in office,” he said, is “perhaps the most abusive and destructive violation of one’s oath of office that I can imagine.” In the wake of Graham’s unexpected death on Saturday, I wrote about him again. He was, in many ways, the quintessential politician of Trump’s Washington, a man who shifted 180 degrees. Graham had been a fervent American patriot, proud reserve officer and military lawyer; he became one of the most important defenders of a president who believes soldiers are “suckers and losers” and shows no respect for law of any kind. As his politics changed, his language grew coarser and his behavior grew worse: He urged Israel to “level the place” in Gaza, and insulted the Danish prime minister in Munich, prompting a colleague, Senator Elissa Slotkin, to issue an apology. Here’s an excerpt from the article, which you can read in The Atlantic (gift link here) Like his close friend John McCain, Graham believed that America should stand at the center of a broad democratic alliance. The few times I met him were in that context, at the Munich Security Conference or at other European-American gatherings, which he once attended with great regularity. He also took seriously the practice of American democracy at home. In a 2014 conversation with The Atlantic, he described himself as a pragmatic politician who eschewed populist slogans. “I know Washington is broken, but what’s broken about it is everybody yelling and nobody trying to fix it,” he said. “I’m trying.” When Donald Trump first appeared in U.S. politics, Graham recognized him immediately for what he was: the spokesman for an alien ideology, one radically different from the idealistic patriotism that Graham had practiced and preached since childhood. Trump privately mocked the military that Graham loved as nothing but “suckers and losers.” He put falsehoods, cynicism, and personal greed at the center of his politics, while expressing disdain for transparency, accountability, and democracy itself. In 2015, Graham described Trump as a “race-baiting, xenophobic, religious bigot” who should “go to hell.” Also as a “nutjob.” When Trump won, Graham understood, as did so many others, that he would have to make some important choices. For a while, he went silent. In the spring of 2016, I saw him at one of those conferences in Europe. He seemed too depressed to speak. But then, like many other Republicans—and, more important, like many other people who have lived under political occupation or experienced radical regime change—he made the decision to abandon his previous ideals, to bury the patriotism that was once so important to him, and to become, instead, a loud, opportunistic collaborator. As I wrote in 2020, it’s very hard to look in to anyone’s heart and to understand their motivations from the outside. But it does seem that Graham had a deep need to be relevant, to be in the game, to be a player, at any cost. And that, rather than any particular achievement, is what he will be remembered for. Continuing to monitor conflicts of interest, ostentatious emoluments, outright corruption and policy changes that will facilitate outright corruption. (Read my original article, Kleptocracy Inc and check out the SNF Agora Institute chart) June 27 Trump diverted funds from more than 900 projects nationwide—including safety upgrades at national parks—to finance beautification projects around Washington, including $700,000 used to fix a White House walkway. June 28 After a years-long Department of Justice investigation into Abbott Laboratories over its management of a baby formula facility where potentially deadly bacteria turned up enough evidence to criminally charge the firm, top department officials instead opted for a financial penalty—illustrating the department’s move away from strict corporate enforcement under the Trump administration. June 29 Defense Secretary Pete Hegseth stacked the Defense Policy Board—which is charged with providing independent national security recommendations to the Office of the Secretary of Defense—with Trump allies, including tech billionaire Marc Andreessen, who has invested in defense contractors like OpenAI and SpaceX, after purging the panel last year. Trump purchased between $1 million and $5 million worth of stock in Axon—which makes about 90% of US tasers—two weeks before ICE solicited bids for a $220 million taser contract. The White House unveiled limited-edition US passports for America’s 250th birthday, featuring a portrait of Trump overlaid on the Declaration of Independence. June 30 Trump made around $1.2 billion from his various crypto companies last year, including more than $500 million from World Liberty Financial and $600 million from sales of his meme coin. Trump is considering issuing 250 pardons to commemorate America’s 250th birthday. These are expected to be largely doled out through the Office of the Pardon Attorney’s informal network of intermediaries. White-collar defense attorneys say they can typically win pardons for clients for $2 million. Last year, Trump administration officials awarded a no-bid contract worth up to $500 million to Clark Construction for the White House ballroom—negotiations in which the president was directly involved. July 1 The Amazon MGM Studios documentary about Melania Trump earned the First Lady more than $10 million. Trump made hundreds of individual stock purchases worth as much as $12.8 million the day before his announcement pausing “Liberation Day” tariffs for 90 days sent the S&P 500 soaring nearly 10%—one of the largest single-day gains ever. Mar-a-Lago and Trump National Doral saw record surges in revenue during Trump’s first year in office, earning $77.5 million and $121.9 million respectively, compared to average annual revenues of $25 million and $75 million during Trump’s first term. Entities across the Persian Gulf paid more than $300 million to Trump’s businesses last year, including $263 million from the sale of half his stake in World Liberty Financial to a firm backed by Sheikh Tahnoon bin Zayed Al Nahyan, a top UAE royal and brother of the country’s president. Kash Patel failed to properly disclose his stock purchase of a Department of Justice contractor, MicroStrategy, worth up to $250,000 last year, despite laws requiring disclosure within 45 days of the trade. July 2 Trump has been directing administration officials to steer more government contracts—so far worth more than $2.4 billion—to Jeff Bezos’s Blue Origin aerospace firm, after Amazon gave $1 million to Trump’s inaugural fund, donated to the White House ballroom project, and reached a $40 million deal to license a Melania Trump documentary. Alongside the billions he made from cryptocurrency last year, Trump earned millions of dollars from real estate licensing fees, including $21 million from organizations in the UAE and $9.2 million in Saudi Arabia, along with smaller sums from merchandise sales, including $4.7 million from Trump Watches and $35,920 from 45 Guitar, a guitar brand endorsed by the president. House Democrats accused consultants tied to the Trump administration of financial fraud, alleging they gave prospective donors to America250—the bipartisan committee created by Congress to celebrate the nation’s semiquincentennial—the banking and routing numbers for Freedom 250, the White House’s fund for events like the UFC cage fight. Democrats also allege that Freedom 250’s CEO traveled to the World Economic Forum in Davos to solicit donations from foreign government officials and business leaders. The Trump administration’s proposed rollback of gun regulations would lift the prohibition on mailing handguns to individuals, a significant win for GrabAGun—the direct-to-consumer firearm retailer on whose board Donald Trump Jr. sits and in which he holds a 1.1% ownership stake. GrabAGun currently must rely on a middleman to transfer firearms to customers. Trump’s financial disclosure revealed that he made more than 21,000 securities trades during his first year back in office, growing his investment account to at least $858 million across stakes in around 1,600 companies, compared to the 13 total trades Joe Biden made during his entire presidency. July 4 Nearly 1 million people who bought the president’s meme coin—Trump Coin—collectively lost $3.8 billion on the token, while he walked away with $636 million. July 6 Shares of Dell closed up 4.4% after Trump recommended people buy the firm’s computers during a speech announcing the launch of Trump investment accounts for children, boosting the more than $1 million worth of shares the president owns in the company. July 9 Federal Reserve Chair Kevin Warsh named crypto investor and Trump ally Marc Andreessen to a task force aimed at modernizing the Federal Reserve, including rethinking its approach to inflation—a key factor keeping interest rates elevated. Palm Beach International Airport has officially changed its name to Trump International Airport—a move that will require Florida to pay royalties to the president’s family for licensing his name. I was there a few months ago, during a trip to Barcelona. The sculptures were a revelation, I’d never seen so many together before. Each one has its own personality. Enjoy the good weather, if you are lucky enough to have it. History Will Judge the Complicit The Quintessential Politician Autocracy Inc Open Letters, from Anne Applebaum is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe
Shaky political science: using "simulated" elections over real-world elections
[Dear DSOS readers – this article is the latest in our “Shaky Political Science” series. In our 38 page report, “Shaky political ‘science’ misses mark on ranked choice voting,” released in early December, we examined over 40 studies focused on ranked choice voting. We found that many of these studies suffered from puzzling research methodologies, poorly constructed surveys, unrealistic simulated elections, cherry picked data, and faulty analyses that often were contradicted by results from real-world elections (a link to our full paper is here). Here are links to the first three articles in our series: the first one summarizing the overall results of our report, the second one analyzing a study co-authored by University of Minnesota academic Larry Jacobs, which was one of the most error-prone of all the research we reviewed in our report, and a third article showing an admirable example of two political scientists holding two other political scientists accountable for their flawed, sloppy research]. In this fourth article in our “Shaky Political Science” series, we shine a spotlight on seven studies that were web-hosted and funded by New America’s Political Reform Program and its Electoral Reform Research Group (the ERRG is an ad hoc committee headed by New America’s Political Reform Program and joined by the American Enterprise Institute, the Unite America Institute and Stanford University’s Center on Democracy, Development, and the Rule of Law). This research consortium has been led by Lee Drutman, senior fellow with New America’s Political Reform Program. These seven studies are simply not credible from any objective standard about how political “science” is supposed to illuminate our understanding of either ranked choice voting or politics in general. The feature that six of these seven studies have in common is that they all used surveys – artificial, simulated elections -- to attempt to say something cogent or interesting about ranked choice voting (RCV) elections. None of them used real world data from actual RCV elections – either the 1000+ elections in the US or the many thousands of elections from places like Australia, Ireland and elsewhere where RCV has been used for 100 years. While a few of these studies conclude with some mixed positives about RCV, our analysis here is content neutral. The idea that, with a political reform like RCV, you can ask respondents in a mock, artificial election to vote with that method, with only the barest of explanation about how RCV works or its likely impacts (i.e. no spoiler candidates, majority winners, voters liberated to pick their favorite candidates, coalition-building), and then ask them their thoughts/reactions about voting in that system, with fake candidates and no real campaigning, no real exposure to an actual election or actual candidates, rests on some very puzzling assumptions. The findings of these seven studies actually are contradicted by the findings from many other studies based on real-world RCV elections. These New America studies also failed to list the real-world RCV studies in their own citations, instead including those studies that reinforced their conclusions. The seven studies featured in this article were chosen because they form a cohesive cohort which precisely illustrates one of the main points of our research paper, namely that using surveys, i.e. artificial elections, or mathematical/theoretical models, instead of real-world data, is in fact “shaky political science.” Below are summaries of our analysis of five of the seven New America studies. You can read our complete analysis of all seven studies at our paper, “Shaky political ‘science’ misses mark on ranked choice voting.” Choosing to “Vote As Usual” by Andre Blais, Carolina Plescia, and Semra Sevi. New America. (Blais et al., 2021). This study, web-hosted and funded by New America’s Political Reform Program and its Electoral Reform Research Group, used two surveys in which respondents completed ballots for simulated elections using plurality, approval, RCV and point (Score) voting. In both cases the authors provided the respondents with only minimal explanation of each method. For example, the entire explanation for RCV was “For this vote, rank the candidates from your first to your last choice (you do not need to rank them all). This is called the ranked vote.” That’s it. The study did not provide any context or explanation to survey respondents about why she/he might want to rank multiple candidates. Participants were not even told that ballots would be tabulated using an “instant” runoff, not to mention any description of RCVs ability to eliminate the spoiler effect or liberate voters to rank their favorite candidates without fear of helping the “greater evil” candidate to win. That’s like taking an iPhone back in time to 1970 and asking someone if they like it better than their rotary phone without explaining how it works. Their first survey targeted states that held 2020 Super Tuesday Democratic primaries and used the real candidates’ names in the survey; the second survey used fictitious candidates. In other words, they surveyed respondents who had only ever participated in plurality elections, had no history or background with using RCV or any of the other methods, had no clue about its likely impacts, and did not communicate how the alternative methods would be counted, and then asked them: “How satisfied are you with using each system?” Predictably, the winner was plurality elections, the system used for president, Congress, gubernatorial and state legislative elections – in other words, the method they have been using, year after year, all their voting lives. With so little explanation of RCV and its impacts, this study is simply not credible in discovering anything useful about real-world RCV elections. It merely shows that a basic voter education campaign is needed when adopting a new election system. Additionally, the result contradicts many other post-election surveys following real RCV elections, asking opinions from voters who had just participated in a RCV election, that showed a great deal of satisfaction with RCV (see details on that in Studies 22-25 in our paper, as well as the four studies cited in the review for Study 28). Share Other studies on “satisfaction” and “understanding” make similar mistakes, in terms of basing their surveys on inadequate explanation to respondents and lack of education about context. How could any survey ever hope to measure accurately, in essence simulate a real-world election, without actual campaigning candidates, the media writing about those candidates, organizations endorsing candidates and voters being inundated with crucial information? More than anything, what these “studies” based on surveys show is that if you ask voters about something they don’t know very much about, and fail to adequately educate them, their responses will unsurprisingly reflect their confusion. Equally puzzling is that New America’s Electoral Reform Research Group, headed by political scientist Lee Drutman, found this to be a credible enough study that they not only published it but highlighted it in Drutman’s own discussion of research related to ranked choice voting. This pattern was repeated in study after study featured by New America’s Electoral Reform Research Group. For example in Ranked-Choice Voting and Political Expression: Voter Guides Narrow the Gap between Informed and Uninformed Citizens by Cheryl Boudreau, Jonathan Colner, and Scott MacKenzie, the researchers based their analysis on a survey which found that respondents who were asked to vote in a simulated RCV election, with minimal explanation of RCV or how it works, displayed less evidence of being informed about how to use RCV effectively, with the least informed voters ranking fewer candidates. As pointed out under the study above, in a real election voters are informed by a wide variety of methods, including public education by the election administration, by interested political groups, the media, and the candidates themselves who benefit from voters being aware of how to vote in the new system. In Support for Ranked-Choice Voting across Race and Partisanship by Joseph Anthony, David C. Kimball, Jack Santucci and Jamil Scott, the researchers used respondents from different parties and races voting in simulated plurality and RCV elections, who indicated their satisfaction with each voting method. Instead of using real world election results, the authors fell back on the easy, armchair and less reliable academic research methodology of making up mock elections. As with the other studies, little explanation was given of how RCV works or what its implications might be (electing majority winners or eliminating spoilers, for example). Unsurprisingly, this study found that respondents were not very satisfied with RCV. However, when the authors added an explanation that RCV has a tendency to elect more women and people of color, they found that this greater explanation increased the preference for RCV. The more voters were familiar with it, and the more they understood potential positive impacts – surprise! -- the more they were open to switching away from the status quo. Of note, the results from this study are thoroughly contradicted by a number of other studies whose data was obtained from real RCV elections with real voters. Those studies show a great deal of satisfaction with RCV. More “fake election” research An interesting wrinkle on this “research by fake elections” approach, which only served to illustrate how unreliable such “research” can become, was taken in Ranked-Choice Voting, Runoff, and Democracy: Insights from Maine and Other U.S. States by Joseph Cerrone and Cynthia McClintock_,_ also web-hosted and funded by New America’s ERRG. It gets high marks for creativity, but an eye-roll for thinking this could possibly add anything meaningful to our understanding of real world RCV elections. This study tried to determine whether RCV, plurality or two-round runoff elections provided greater voter satisfaction and acceptance of election results. To determine that, the authors conducted a survey asking if voters were satisfied with a fictional newspaper report of an election, with three different versions given to respondents – plurality, runoff and RCV, with variations of Republican/Democrat winners and variations of “come-from-behind” winners,” i.e. candidates not initially in first place but who win the “runoff.” Again, a minimal explanation of RCV or how it works was given, and the survey showed that respondents had the greatest unfamiliarity with RCV than the other methods. The unsurprising result? Respondents were much more satisfied with “come-from-behind” victories in a traditional two-round runoff election than in RCV, even though RCV is also a runoff system but with a single round of counting. Interestingly, the dissatisfaction with RCV disappeared for those who were more familiar with RCV. In fact, among those respondents who were “very familiar” with RCV, they were equally satisfied with RCV, runoff and plurality. Surprise, surprise. Again, what this study revealed, more than anything, is not only the inadequacy of the explanation for RCV given by the authors of this study, but the failure of the research format itself, thinking that such an artificial mock election could somehow simulate what actually happens in the real world in a real RCV election. Those four studies reviewed above are fitting examples of inadequate and shaky political “science.” Other studies reviewed in our paper from the New America/Electoral Reform Research Group cohort which also used mock elections included: Ranked-Choice Voting is an Acquired Taste by Joseph Anthony and David C. Kimball. New America. (Anthony et al., 2021). Ranked-Choice Voting is No Refuge for Extreme Candidates by Melissa Baker. New America. (Baker, 2021). Also included in this New America cohort is Does Ranked-Choice Voting Reduce Racial Polarization? by Yuki Atsusaka and Theodore Landsman, which actually did use real-world election results data but applied an odd and unconventional methodology to the data which relied on shaky assumptions and non-real-world scenarios that were too divorced from reality. To determine levels of “racially polarized voting,” the researchers imagined “ethnic parties,” by which they mean political parties for each ethnic group, which was especially odd since the election data they were using came from nonpartisan RCV races. But then they made an even odder assumption, namely that since race and ethnic characteristics of studied voters were not available, the authors tried to infer that information based on who voters selected and the racial make-up of voters in each district. This amounted to little more than vague and ambiguous guesswork dressed up with a bunch of mathematical equations used to give it a veneer of scientific appearance. Indeed, it seemed to assume racial polarization in order to prove racial polarization. Based on this questionable methodology, the authors concluded that RCV doesn’t reduce racial polarization. But the premise of this study missed an obvious reality: since RCV incentivizes more candidates to run and can result in more diverse fields of candidates, one would expect more racially-polarized voting in voters’ top rankings, not less, because more voters would have someone on that ballot who represents their racial or ethnic identity and more freedom to vote for their real first choice. By encouraging more diverse candidates and eliminating the spoiler effect, RCV facilitates voting for a favorite candidate of your own race and then selecting a second choice with the best likelihood of winning. The real measure is what voters do with their additional rankings – do they consider other candidates and vote “for the best of the rest”? This study didn’t tell us anything useful about that. But readily available data shows this is common in real RCV elections, and candidates act on this knowledge by engaging with voters outside their racial group. So the authors of this study failed to consider or account for what actually happens in real-world elections, and their conclusions are tainted by this failure. New America has contributed a great deal over the past 2.5 decades to our understanding of many policy areas, including politics and political reform (full disclosure: Steven Hill, one of the authors of this article and of the “Shaky Political Science” research paper and series, is a former senior fellow and former director of New America’s political reform program, 2005-2010). So it was especially disappointing that New America funded and published such a large batch of dubious studies about ranked choice voting. The stated praiseworthy goal of its Electoral Reform Research Group was to inspire new academic research about ranked choice voting that would shed light and help the public better understand the impacts of RCV. Instead, the ERRG research did just the opposite. It produced dubious research that created more confusion and misinformation. In so doing, New America’s political reform program contributed to “shaky political science.” Steven Hill and Paul Haughey Share Upgrade to $5 subscription Thanks for reading DemocracySOS, a reader-supported digital portal for the pro-democracy movement. Subscribe for only $5 per month to receive full benefits and to support our work. Subscribe
“AI polls” are fake polls
A few weeks after Donald Trump’s second presidential win, I took the train up from London (where I was living at the time) to Oxford to attend a conference on polls and forecasts of the 2024 election. Most of the attendees were pollsters or academics, but I also watched presentations from Aaru and Electric Twin, two companies that do what is interchangeably called synthetic sampling, silicon sampling, or synthetic audiences. Stripped of startup jargon, that means they use large language models (LLMs) to simulate responses to public opinion polls by having AI agents take on the role of survey respondents. I had already heard of Aaru thanks to some articles with eye-catching headlines like “No people, no problem: AI chatbots predict elections better than humans” in the months leading up to Election Day. The founders were making some big, some might even say far-fetched claims, such as: “within two years, we will simulate the entire globe — from the way crops are grown in Ukraine to how that impacts production of oil in Iraq, trade through the strait of Malacca, and elections for the mayor of Baltimore.” When Semafor asked Aaru’s cofounders — Cameron Fink and Ned Koh — about my boss, they said “we respect all those who came before us.” Nate (as he so often does) shared his thoughts on Twitter: Fink and Koh were relatively good-natured about this back-and-forth when we spoke at Oxford. They even offered to mail me one of the t-shirts featuring Nate’s quote they apparently had made. I never took them up on the offer, which I now somewhat regret. These synthetic sampling companies fell off my radar for a while, but they do still exist. In fact, Aaru recently received a $1 billion valuation. Is what they’re doing anywhere close to the most important frontier in AI development? Not by a longshot, especially when Anthropic just developed a model so adept at exploiting software vulnerabilities that it’s only being released to 40 companies. Still, silicon sampling is increasingly finding its way into public polling. Axios reported in March that “a majority of people trust their own doctors and nurses” based on findings from Aaru — without mentioning that the “people” in that sentence were actually LLMs. Around the same time, the Public Sentiment Institute “boosted” their online sample of 373 real survey respondents with 114 AI agents.1 (Spoiler alert: even the co-founder of Electric Twin doesn’t think that’s a particularly defensible approach.) Polling companies like Qualtrics and Ipsos are also developing synthetic data panels. So, what should we make of these … “polls”? Let’s get one thing out of the way: whatever they are, they’re not polls in the way that term is usually defined. Subscribe You can’t replace polls with AI On one hand, using LLMs to essentially make up fake survey respondents sort of sounds like the dumbest idea ever, one that will at best imperfectly replicate real polls while introducing all sorts of biases. On the other hand, with LLMs improving at a remarkable, perhaps even alarming rate, maybe that means I’m a dinosaur at the ripe old age of 24 because I still want to rely on polls that talk to actual people. I’m not going to argue that synthetic samples are completely useless. In fact, as I’ll return to later, there is evidence that some techniques can replicate topline survey results quickly and cheaply. But the marketing from certain companies can be slightly optimistic. “No traditional poll will exist by the time the next general election occurs,” said Fink in 2024. We’re just 206 days away from the midterms, and based on the fact that I still have to collect a bunch of polls every day, I’d say he should have run that prediction by a sample of AI agents before the interview.2 To see why synthetic samples can’t replace polls, here’s a quick primer on how they work. The simplest version of these models involves taking a LLM (like ChatGPT or Claude), giving it a demographic profile (e.g., a white, college-educated woman who lives in Utah and makes $70k a year), and then asking it to respond to a survey question. You repeat that process a few thousand times using different demographic profiles and end up with a sample of synthetic survey responses. The actual models used by private companies are more sophisticated than this, usually because they incorporate more hypothesized demographic characteristics for each agent and provide them with extra information. Aaru, for example, feeds agents a diet of news and information they’d be likely to consume, while Electric Twin incorporates their customers’ proprietary data about the audience they’re trying to replicate. The way Ben Warner, the co-founder of Electric Twin, explained it to me was “we have a large amount of data on […] for instance, 5,000 people. Can we make an accurate prediction of how they would respond to another question?” Still, it should be obvious why synthetic samples can’t replace polls. Polling is fundamentally a data collection process. We might use surveys to make predictions by feeding them into election forecasts, but the main purpose of a poll isn’t prediction, it’s gathering new data about what people think and how they feel. Silicon sampling, on the other hand, produces no new data. It’s simply a model: you input LLM training data, demographic prompts, and a bunch of other information, and it spits out a prediction for what a poll would say. We love models here, but models aren’t polls. That difference is an important philosophical sticking point for most pollsters I talk to. “I think politics should stay away from [synthetic sampling], because we’re trying to […] represent the voice of the people,” said Natalie Jackson, a vice president at GQR Insights. Democratic pollster John Hagner told me: “I think I’m just incredibly skeptical of this idea. I don’t think it’s research. At that point, you’re asking the machine to tell you what you already believe.” Hagner has seen some presentations of early synthetic sampling experiments, but so far, “if it’s being used in a campaign, people are keeping it incredibly quiet.”3 But Eli, I hear you saying, aren’t polls themselves increasingly governed by modeling decisions? Indeed they are: pollsters’ choices on which sampling method to use, how to define their likely voter models, and how to weight their samples can and do lead to dramatic differences in the results they publish. Aaru even referenced these limitations in the methodology statement included with that maternal mortality “poll” — although I’m using the term “methodology statement” loosely here, because it doesn’t really explain how the model works at all. We can ignore the (frankly preposterous) implication that synthetic sampling isn’t subject to a separate set of biases. The important point is that there’s still a meaningful difference between using weighting and other statistical techniques on actual polling data and using a model to predict what a poll would say. The latter is far closer to election forecasts or techniques like MRP — potentially useful models, but not a replacement for polls.4 To be fair, other synthetic sampling companies are perfectly happy with the distinction between polls and models. Warner compared polling and synthetic sampling to different tools in a toolbox. “The mistake I think we make is we think that these new tools should either work in exactly the same way or somehow replace these old tools,” he said. “Rather than thinking of it as, okay, so we’ve always had the hammer, we’ve always had the screwdriver, now we’ve got a saw. But don’t use a saw to try [to] do the job of a hammer.” A quick comment from Nate Eli didn’t ask me for a comment — rather rude of him, don’t you think? But since I’m editing this story, I figured I’d add a few quick thoughts rather than putting words in his mouth. Beyond the frequently misleading marketing, what bothers me about the AI “poll” hype is that as AI tools make statistical inference cheaper and/or better (note that these are not synonyms) that actually increases the comparative value of collecting original data. You might be able to train a model to make a reasonable estimate of what some hard-to-reach poll respondent would say — say, a young Black man who voted for Trump. (Such a person checks a number of boxes for a voter who is usually hard to reach in surveys.) Indeed, this is closely related to what models like the Silver Bulletin forecast already do. They essentially smooth out the kinks in noisy survey data by making inferences based on past voting patterns or national polls or surveys of other states. But you don’t actually know what these voters think unless you’re reaching them directly. If there’s a shift in opinion among this subgroup, you’re not going to detect it. So if I were running a campaign, I’d invest more in going the extra mile to find a representative sample of those voters. And then I’d hire some smart quants — or Claude? — to figure out the implications for campaign strategy based on proprietary data that my competitors didn’t have access to. -Nate Silver Are these models any good? If synthetic surveys are just a new type of model, the next obvious question is whether the models are at least accurate. The answer very much depends on who you ask. On one end of the spectrum, you have the maximalist argument that synthetic sampling is more accurate than actual polls. “It’s an incredibly challenging problem to go to someone and say ‘hey, we’re going to be more accurate at predicting human behavior than you, even when you talk to your customers directly’,” Koh recently told CNBC. In his view, synthetic sampling isn’t a saw to polling’s hammer, it’s “magic.” There’s certainly evidence that synthetic samples can replicate certain survey toplines. But if Aaru does have any examples of their approach outperforming the polls, they’re keeping those to themselves.5 Aaru’s 2024 election model, for example, had Kamala Harris leading in Michigan, Nevada, Pennsylvania, and Wisconsin on November 4th. And although they’ve since taken down their forecast page, they gave Harris a 50.5 percent chance of winning the race on November 2nd.6 After the election, Fink told Semafor he was happy enough with those results because they were “within margin of error,” a term that is completely meaningless when applied to a “sample” of AI agents. And of course, Aaru says their models have improved since 2024, so supposedly now they’d be more accurate than the polls? Still, their stronger argument is on cost: “We are significantly faster and cheaper than traditional polling, and still more accurate,” said Fink. The first two claims are undeniably true, but the third brings us to the opposite end of the spectrum. Both Jackson and Hagner are skeptical that these models are reliable for anything beyond replicating common survey toplines. “I just […] don’t think the machines are what we want when we’re looking for nuanced views. My example on this is people in Arizona and Nevada in 2024 who voted for Trump and voted for expanding abortion in their states on ballot initiatives,” said Jackson. Hagner identified another issue. Maybe the synthetic respondents, like sycophantic LLMs, are inhuman in one important way: they’re too nice. “The reports that have come through at the meetings that I’ve been at are that the early experiments on this, they cannot get respondents to be as racist or sexist or, frankly, as negative as human respondents,” he said. Academic research mostly agrees on this point. While there are some papers that show promising results when using LLMs to replicate polling data, most show that LLMs suffer from various quirks like producing too few “don’t know” responses and can seriously overpredict the favorability of politicians like Donald Trump and Kamala Harris. They also seem to struggle with too little variation between demographic subgroups, so the difference in predicted opinion between Democrats and Republicans, for example, is too small. When I asked Warner about these studies, his response to these papers was that just because academics can’t get synthetic sampling to work doesn’t mean that the technique doesn’t work in general. “Actually, the argument is, okay, yours does not [work]. That does not mean […] for this complex set of machinery, which uses a lot of investment, a lot of time, a lot of money, you can’t get it to work.” Cards on the table, I’m somewhat sympathetic to this argument because academics aren’t exactly great at making election forecasts. Usually, the people with skin in the game are the most accurate. Warner’s argument is that the approach Electric Twin takes — which includes, for example, making multiple predictions for each synthetic respondent using different models and prompts and subsequently averaging those to get a final prediction in a sort of ensemble forecast — produces better results than the simpler academic models. Warner shared a comparison between his method and the method from a recent academic paper with me, and Electric Twin was indeed able to get more accurate replication. But even still, he acknowledged that synthetic sampling “is not a crystal ball.” “If you asked me, do I think using other data sources will be more accurate than asking somebody who they will vote for, I would probably say no. But if you asked me ‘would your system be useful for our turnout modeling today?’ I would say yes.” For better or worse, it looks like the method is already getting more popular in the market research world. Most of the clients Aaru touts these days are businesses like EY and McDonald’s. And AI will probably begin to pop up in other parts of the political polling process. Pollsters are already using it to code open-ended survey responses, and some firms, like YouGov, are testing using LLMs to ask survey respondents questions. More worryingly, one danger to actual polls is that AI agents can be used to infiltrate online surveys. Most online polls use various checks to prevent that from happening, but there’s conflicting evidence on how effective those filters are and how prevalent AI agents currently are in online panels. If those agents ever become impossible to detect, it might spell the end of online polling, but the solution isn’t to replace all of your respondents with ChatGPT. Silver Bulletin is a reader-supported publication. To receive new posts and support our work, consider becoming a subscriber. Subscribe 1 That particular poll obviously doesn’t meet Silver Bulletin standards for aggregation. But we exclude all Public Sentiment Institute polls from our averages because we classify them as an amateur polling firm. 2 You could argue that Fink meant the next presidential election, but (a) I’m also confident we’ll still have real polls in 2028 and (b) in that case he should have asked an LLM to define “general election.” 3 Quick caveat: that’s reporting from a Democratic pollster. It’s possible that Republicans are more willing to use AI in political campaigns. 4 Indeed, Silver Bulletin does not include “polls” produced by MRP in our forecasts or averages, and we think it’s extremely misleading when their practitioners describe them in a way that suggests original data had been collected among a large number of states or Congressional districts. 5 A recent report from Aaru and EY did show two examples of a synthetic estimate being closer than a survey to a real-world benchmark — but I’d take those findings with a grain of salt because the report reads more like an ad and didn’t involve any sort of prediction being made ahead of time. 6 For comparison, our odds for Harris on the same day were 48.2 percent.
Your Thymus and Your Healthspan
Our thymus gland plays a central role in the development or our immune system, specifically for supporting T cell development (how these cells got their name, maturation and differentiation in the thymus) and discriminating between self and foreign, non-self antigen proteins with production of dendritic cells. As we age, our thymus gland shrinks—the process known as involution—with progressive change from spongy to fatty tissue, with loss of functionality. But this process, with respect to timeline and extent, markedly varies from one person to the next. And to make things even more complicated, our thymus gland anatomy and precise location in the chest also is highly variable. So the famous Frank Netter diagrams from the 1960s (such as at left below) don’t capture the remarkable heterogeneity that a random set of 12 chest CT scans (below at right). What I consider as 2 landmark papers in Nature this week (here and here) used AI to quantify health of the thymus in 2 large cohorts, and then correlate the metric to a broad array of health outcomes. In this edition of Ground Truths, I’m going to cover three questions: (1) How was thymus health determined by AI?; (2) How is thymus health linked to key clinical outcomes? and (3) Are there ways we can promote a healthy (aka rejuvenate) thymus gland in our later years? Subscribe 1. How was thymus health determined by AI? This was accomplished by extensive work developing and validating an AI pipeline for 2 stages, the first for localization and segmentation of the thymus bed, and the second to quantify a digital marker, termed the thymic health score. There’s a lot to this algorithmic development, so I won’t go into all the details, but just provide a rudimentary outline of what was done. The 1st stage used supervised learning with 2 radiologists reviewing 2,461 chest CT scans. That work trained a 3D U-Net model to automatically identify the thymic bed’s 3D cropped region using center of mass coordinates, which reached 99.8% accuracy. The 2nd stage used a foundation model that was pre-trained with SwAV (which stands for Swapping Assignment between multiple Views), a self-supervised model, that generated a high-dimensional thymus representation of 4,096 features. That compares to the radiology reductionist and subjective scoring of 0 to 3 where 0 means the thymus is fully degenerated and fatty, and 3 is considered dense, intact glandular soft tissue. Notably, the self-supervised learning (SSL) performance was superior to supervised (AUC ~0.75 vs 0.55, respectively) reflecting the shallowness of the 0-3 classifier vs SSL’s holistic, self-taught AI. The thymic health score ranged from 0 to 100, with the highest number indexed to fully preserved thymic health. The algorithmic work extended to explainability with both Shapely value distribution, meaning the model was holistic not relying on any “magic” pixel, and occlusion sensitivity, demonstrating the model’s performance was not affected by ribs, sternum, lung tissue, but rather focused on the thymic bed directly. Below are a few saliency maps from the occlusion sensitivity that indicate the thymic bed focus, not affected by neighboring structures, using the jet color scale. Now the AI was ready for processing ~25,000 CT scans, each over 30 MB files, from the 2 large participant cohorts to provide a thymic health score, it did so in less than 14 hours, which is less than 2 seconds per scan! Share 2. How is thymus health linked to key clinical outcomes? The two cohorts were the National Lung Screening Trial, NLST (N=25,031) and the Framingham Heart Study, FHS (N=2,581) each with baseline demographics and long term follow up of >12 years for health outcomes. From the NSLT, you can see below that increased age and body-mass index were correlated to reduced thymic health. And, overall, men had reduced thymic health scores compared with women. There are a lot of graphs in the paper for outcomes, for both the NLST and FHS cohorts. To simplify the major outcomes, below is a composite Figure with striking reduction of all-cause mortality and for the different causes of death. The results were consistent between the 2 cohorts. For example, the cardiovascular mortality hazard ratio average was 0.57 in NLST and 0.38 in FHS (weighted to be 0.53 for the aggregate). The reduction of mortality also extended to digestive diseases and pulmonary disease (data not shown below). Lung cancer incidence was reduced by 36% for high vs low thymic health (Figure below, with a panel -incidence and b panel-mortality, adjusted for age, sex, BMI, smoking status). Notably, there was a smoker’s paradox: lower incidence and improved survival for lung cancer in smokers who had a high thymus health score. Smoking was associated with lower thymic scores whereas alcohol was not. Higher thymic health was linked to higher HDL cholesterol, lower triglycerides, lower fasting blood glucose, and lower systolic and diastolic blood pressure. Inflammation biomarkers, such as C-reactive protein and interleukin-6, were elevated in the people with lower thymic health scores. In their second paper, the same team looked at the relationship between thymic score and outcomes for 3,476 patients receiving cancer immunotherapy. The non-small cell lung cancer progression-free survival was reduced in patients with low thymic health score. A parallel relationship for survival was seen for other types of cancer including melanoma, breast, and kidney with a 44% lower risk among high thymic scores (Figure). A similar pattern was seen with different immune checkpoint inhibitors. The thymus health score outperformed programmed death ligand-1 (PD-L1)and tumor mutation burden (TMB) assay of the tumor and was an independent prognostic marker of progression-free survival. In this dataset, a direct connect with thymus score and adaptive immune function was noted with a correlation before cancer treatment for both T-cell diversity and thymus T-cell production (a metric known as T cell receptor excision circles). Thymus Removal Study There is a highly relevant citation about thymus function before moving onto ways to rejuvenate it. In 2023, a very important study of thymus removal during cardiothoracic surgery was published, with 1146 patients undergoing thymus removal vs 1146 matched controls (mean age 55 years). The adverse outcomes for thymus removal were striking, with a 2.9 higher risk of all-cause mortality and doubling of cancer incidence, 1.5 fold increase in cancer mortality, and 1.5-fold increase in autoimmune diseases. The immune function showed marked compromise after thymus removal, as reflected by CD4+ and CD8+ T cells (Figure below) Share Ground Truths 3. Are there ways we can promote a healthy (aka rejuvenate) thymus gland in our later years? In recent months we’ve learned a lot about the process of thymus involution and the pathways by which this may be modulated. Thymus involution is primarily due to loss to thymus epithelial cells (TECs) and there are 2 subtypes with different functions: the cortical, responsible for positive selection of T cells, and the medullary, for negative selection. When adipose tissue infiltrates the thymus, thymic adipocytes are pro-inflammatory, knocking out T-cell output. Likewise, systemic inflammation, or “inflammaging,” from smoking, obesity, and chronic stress promote thymic adipocytes. Now, from preclinical studies, we know about key pathways that account for these processes. Historically, back in 2014, FOXN1 was recognized as a single master transcription factor essential for TEC function, with studies in mice with forced FOXN1 upregulation leading to thymus regeneration. That laid the foundation for potential regrowth of quiescent thymus cells. It took awhile to find pathways to accomplish this goal and to even refute its role as the master regulator. The Liver-Thymus Axis Feng Zhang and his team recently showed that injection of an mRNA vaccine to the liver encoding DLL1, IL-7 and FLT3-L in combination, factors known to support thymopoiesis (Figure), successfully enhanced immune function in aged mice. However, the effect was transient. The FGF21 Story Two reports in 2025 addressed the thymus gland’s production of FGF21, a growth factor (fibroblast growth factor 21) and its control of thymus involution. This growth factor is also produced by the liver, but it doesn’t have impact on the thymus in aged rodent models. In contrast, The liver’s hepatocyte growth factor (HGF) has been shown to reverse senescent TEC structural changes. Growth hormone stimulates production of insulin growth factor-1 (IGF-1) in the liver and the thymus. Two small studies in human participants (TRIIM, for Thymus Regeneration, Immunorestoration, and Insulin Mitigation) of 6 and 50 people, respectively, treated with combined human growth hormone, metformin and DHEA, suggested the potential of slowing epigenetic aging and improving thymus mass and function. But the combination of drugs, small sample, and lack of controls make conclusions murky. Increased thymus FGF21 led to increased CD8 T cells in old mice and extended healthspan and improved physical performance. Building on the previous FOXN1 data, ablation of β-klotho, the obligatory co-receptor for FOXN1, accelerated thymic aging. The benefit to FGF21 was confirmed in a companion paper using an FGF21 knock-in mouse model showing thymus enlargement and increased TECs throughout the lifespan. (schematic Figure below). The accompanying editorial for these 2 papers speculated: “The proposed link between age-associated thymus involution and organismal aging is intriguing, and suggests we might seek to achieve systemic rejuvenation by counteracting thymic involution.” Other Factors Lessons from the Axolotl (Mexican Salamander) In an elegant set of experiments of the axolotl (Figure), known for its complete thymus regeneration after total surgical removal, mediators were identified. Surprisingly FOXN1 was dispensable. But midkine, a growth factor, appeared to be the driver, the initiator of thymus cell growth, survival, and repair. Other components that contributed were Postn+ (periostin-expressing mesenchymal niches) and Ccl19, a chemokine protein from the bloodstream. RANK and RANKL RANK (receptor activator of nuclear factor κB its ligand (RANKL) have been considered key regulators of medullary TECs. In a study of both aged mice and human thymus cell cultures, their key role was confirmed, with multiple benefits of resorting TEC function, recruitment of progenitor cells, and T cell development. Subscribe Summing Up This body of work strongly supports the health of the thymus gland as a critical regulator of human healthspan, not just a correlate or link. The new landmark studies are reinforced by the regenerative biology experimental models that have shown, via an array of mediators, that healthspan of aged mice is at least in part dependent on thymus gland function and its critical role for maintaining adaptive immunity. It sure seems that we’d be better off having an immune reservoir of naive T cells instead of a profile in older adults of exhausted, low quality, memory T cells, with a senescent phenotype. It is striking that we ignored the importance of the thymus gland for many decades. A wake-up call was the results of the thymectomy matched control study reviewed here, with the opposite of high thymic health score improved health outcomes. The convergence of AI and healthspan here is notable. The exceptional and laborious work for developing and validating (no less explaining) the AI to quantify thymus health set the foundation for probing the connect the major health outcomes in two cohorts with extensive follow-up. We’ve never had a way to meaningfully quantify the thymus gland before, so this represents important leverage of supervised, self-supervised, and transformer models. The study wouldn’t have been possible without AI. And it begs the question as to whether there are ways we can maintain our thymus at a high health score. But it’s not so simple to rejuvenate our thymus. There are multiple risks including the induction of autoimmunity, increasing the risk of cancer, and inducing a pro-inflammatory state. For example, factors that increase TEC proliferation could compromise the medulllary negative selection process and leak out auto-reactive (self-attacking) T cells. Thymic cells could be induced to proliferate, but we know, for example, that growth hormone and IGF-1 are associated with an increase risk of cancer. If thymus rejuvenation doesn’t clear out the accumulated senescent cells in older adults, then inflammation within the gland could block the production of TECs. Keeping these risks in mind, we now have identified many ways to keep our thymus healthy as we age that undoubtedly will be tested in clinical trials going forward. Ultimately, the benefit-to-risk tradeoffs will be defined. In the meantime, we have desperately needed an immunome. As I’ve stressed in multiple prior Ground Truths, we have no test in the clinic to assess a patient’s immune status! It was very encouraging for me to get an email from Hugo Aerts, the senior author of the landmark papers in the days after they were published: “I also wanted to say that your Substack on “Why We Need an Immunome” was inspiring for this work. We found it a very compelling perspective and greatly enjoyed reading it.” Just having an assessment of a person’s immune status via a chest CT scan, with millions of these performed per year, might help in guiding the right immunotherapy for cancer (e.g. more intensive, combinations for those with low thymic health scores) and help define people who are increased risk for age-related diseases of cancer, cardiovascular, and neurodegenerative. The main theme of Super Agers is that we now have many layers of data (genetics, proteins, biomarkers, organ clocks that can be integrated via multimodal AI) for us to define a person’s specific risk decades before one of these age-related diseases leads to symptoms, enabling prevention. From the new studies this week, we can add thymus health score as yet another way to understand a person’s risk, a window into their adaptive immune system and healthspan. And perhaps someday we will be able to safely rejuvenate the thymus, or maintain its health through one’s life, and use the thymus health score to monitor. NB: I wrote this post (No A.I.) A Quick Poll Loading... ********************************************************************* Thanks to Ground Truths subscribers (now > 200,000) from every US state and 212 countries. Your subscription to these free essays and podcasts makes my work in putting them together worthwhile. Please join! If you found this interesting PLEASE share it! Share Ground Truths Paid subscriptions are voluntary and all proceeds from them go to support Scripps Research. They do allow for posting comments and questions, which I do my best to respond to. Please don’t hesitate to post comments and give me feedback. Let me know topics that you would like to see covered. Leave a comment Many thanks to those who have contributed—they have greatly helped fund our summer internship programs for the past two years. It enabled us to accept and support 47 summer interns in 2025! We aim to accept even more of the several thousand who will apply for summer 2026.