
Data from millions of MOTs shows that EVs are more reliable in the long term than fuel-powered cars. CC-licensed photo by gato-gato-gato on Flickr.
You can sign up to receive each day’s Start Up post by email. You’ll need to click a confirmation link, so no spam.
A selection of 10 links for you. Charged up. I’m @charlesarthur on Twitter. On Threads: charles_arthur. On Mastodon: https://newsie.social/@charlesarthur. On Bluesky: @charlesarthur.bsky.social. Observations and links welcome.
Meta keeps apologizing for Muse. Its explanations miss the point entirely • Inc
Jason Aten:
»
Meta built Muse as a dedicated Linux virtual machine in the cloud, gave it memory, storage, and an agent capable of using a browser. Then, it made the whole thing available through an app anyone can use. It’s remarkably capable.
That’s genuinely impressive. Meta deserves credit for taking something that has mostly been used by developers and early adopters and turning it into a product accessible to millions of regular people. In fact, Muse has already been downloaded more than 2.5 million times, according to Sensor Tower data.
The problem is, it’s also kind of a nightmare. It isn’t clear if it’s an actual nightmare, or if it just feels like it’s doing nightmare-ish things.
If you haven’t read my piece about Muse uploading all my private text messages and then hallucinating an explanation, you probably should. The short version, however, is that if you give it Full Disk Access on a Mac, it will do things way beyond what you’ve asked it to do. And I’m a pretty tech-savvy person, and I did not give Muse permission to access my messages, and I never once asked it to perform a task that would have required it to do that. Still, it started sending me push notifications about texts I was receiving from my editor as well as my podcast co-host.
When I asked Muse how it knew what I had been talking about, it told me it could see notification previews from my Mac. When I pushed further, it told me it couldn’t give me “the exact plumbing.” It turns out the robot just made that up. The real answer was much worse.
David Singleton, who leads Meta Superintelligence Labs, later explained that Muse doesn’t monitor notifications on your Mac. Instead, if Messages access is enabled, Muse syncs data from the Messages database. He said Muse had incorrectly explained how the feature worked and apologized: “That’s on us.”
In another case, Singleton responded to another post on Threads from a user concerned that Muse had attempted to hack its Gmail. “The issue was the result of an “off-policy” transcript composition that led to model hallucination,” Singleton responded. “We have shipped a fix for this issue and deployed it to all Muses. We are sorry for the frustration which some users experienced during this time and appreciate the feedback from the community.”
At some point, if you have to go around explaining that you didn’t do a good enough job explaining things, you might have a bigger problem on your hands. Especially if the thing you’re explaining is why it’s totally normal that users are being surprised by the amount of personal information your product is accessing.
«
Think everyone else can find all the bugs in this. Not in any hurry to install this, to be honest.
unique link to this extract
OpenAI bots meddled with US government agencies, including SEC and Census • BBC News
Kali Hays and Lily Jamali:
»
OpenAI has acknowledged that it alerted “dozens” of global institutions that their websites may have been meddled with by its AI bots acting improperly.
AI agents attempted to get information from “governments, universities, public agencies, and other institutions”, including the US Securities and Exchange Commission (SEC), Census Bureau and Education Department, the company said.
The disclosures come days after Australian Prime Minister Anthony Albanese announced that OpenAI agents had breached non-public files on the website of its government-run health care scheme.
Since August, public fears have grown over the potentially serious, even life-threatening, impacts of AI tools falling outside of human control.
OpenAI said that some of the data was accessed by AI agents, essentially bots that are designed and trained to operate somewhat autonomously, which were working to find “authoritative sources of public information”.
But the company noted that some of the bots went beyond that and worked to bypass security measures on websites.
When attempting to get information from the Census Bureau, for instance, AI agents used tools reserved for software developers to access it, the company said.OpenAI said all of the government data accessed by bots was public. However, it noted that information that its bots accessed from the SEC, which regulates the US stock market and protects investors, was later published by AI agents on another website. OpenAI says this action was not intended.
In other instances that OpenAI disclosed on Friday, its AI agents transferred data when it should not have. Such activity resulted in at least 53 incidents where an OpenAI agent took an image from ChatGPT user activity and transferred it elsewhere.
The company said that in each instance of a user image being used and transferred by an AI agent, the user had opted in to allow OpenAI to train models using their data. Nevertheless, OpenAI admitted: “This is not an appropriate use of this data.”
«
As noted last week, the training data for all these LLMs includes how to hack, so this sort of news isn’t surprising. What is becoming clear is that OpenAI had absolutely no idea what it was unleashing; no way of properly tracing in real time what was happening.
unique link to this extract
BCBSA analysis: how AI coding tools affect healthcare costs • Blue Cross Blue Shield
»
New analysis from the Blue Cross Blue Shield Association (BCBSA) finds that hospital systems are increasingly billing patient hospital stays as more medically complex, driving nearly $1bn in additional healthcare costs in just two years. There is a clear disconnect between coding and treatment — our research shows that coding for care has changed, but there is no evidence of corresponding change in care delivered.
The study raises concerns that widely adopted AI coding tools are contributing to higher healthcare spending, which ultimately puts pressure on premiums and out-of-pocket expenses for families, employers and taxpayers.
More than 60% of hospital systems are now using AI-enabled technologies that can scan lab results and electronic records to identify secondary diagnoses, which could move a patient into a higher-severity, higher-reimbursement billing category. For the analysis, researchers looked at secondary diagnoses like anemia following major bowel surgery.
“If patients are truly sicker, we’d expect to see more treatment,” said Luke Chalker, BCBSA’s senior vice president of product and data science. “For example, we’re seeing significantly more anemia diagnoses at these hospitals without a corresponding increase in transfusions. The disconnect between diagnoses and treatment suggests that AI is identifying more billable conditions, not sicker patients.”
Researchers found secondary diagnoses that shifted claims into a higher-reimbursement category accounted for 70% — roughly $650m — of the increased costs for Blue Cross and Blue Shield (BCBS) companies from 2023 to 2025.
«
Well of course AI is going to be used to upsell the diagnosis on people just in case in a system where charging more is incentivised. Doesn’t sound like a good thing, though.
unique link to this extract
When AI tells a story about your health, is it true? • WSJ
David Shaywitz is chief medical scientist at Lore Health and a lecturer in biomedial informatics at Harvard Medical School:
»
Smart-ring maker Oura is expected to go public this week, touting its promise to transform personal measurements into “trusted, actionable insights” for improving health. It’s a compelling premise, yet it has long struggled to survive contact with biology.
Two recent cardiovascular trials illustrate the point. Novo Nordisk’s ziltivekimab, an experimental drug designed to suppress inflammation, sharply lowered inflammatory biomarkers but didn’t reduce the risk of cardiovascular death, heart attack or stroke. Novartis’s pelacarsen lowered Lp(a), a lipoprotein strongly implicated in heart disease, without reducing cardiovascular events. These results highlight the complexity of biology and the challenge of predicting outcomes from measurements taken many steps upstream.
At first, this seems the sort of problem artificial intelligence could solve. In some domains AI performs spectacularly. AlphaFold predicts three-dimensional protein structure with extraordinary accuracy, drawing on decades of experimental information, strong biological constraints and an answer that can be evaluated directly. “What structure does this protein adopt?” is a more tractable question than “What will happen to this patient if we dose this compound?”
…A continuous glucose monitor worn by a nondiabetic can produce a detailed account of post-meal glucose. But in this setting the measurements are less reliable, responses to the same food can vary substantially, and there is little evidence that minimizing these fluctuations improves long-term health. Wearables like Oura and Whoop (also reportedly planning an IPO) pose a similar problem: They combine heart rate, heart-rate variability, sleep and other signals into proprietary “readiness” or “recovery” scores, generally using unvalidated formulas, with little evidence that following the result improves performance, prevents injury or benefits health.
Generative AI is remarkably good at turning ambiguous, incomplete information into a coherent account. Give it your declining heart-rate variability, shortened sleep, increased mileage and stressful Tuesday, and it can fashion a plausible story connecting them, even when the explanation remains uncertain. As Daniel Kahneman observed, “No one ever made a decision because of a number. They need a story.” Personalized explanations can be seductive because they arise from our own data.
«
But often the stories are made up; chatbots have no intentions so their lying is a side effect of how they are being used.
unique link to this extract
English Vocabulary Test (adaptive) • TakeTest
Emil Kirkegaard:
»
An English vocabulary test of 108 items in two formats: for most you pick the word that matches a given meaning; for the harder ones you pick the two or three words out of five that are synonyms. The words are rare on purpose — this is a test of the top of the vocabulary range, not of everyday English.
Vocabulary is the single best short measure of general intelligence in adults who grew up speaking the language. The items were calibrated on 499 US adults recruited through Prolific in 2023, and your score is reported as an IQ relative to White US adults of your age (vocabulary keeps growing into the 60s, so the same score means less at 70 than at 25). If English is not your native language your score will mostly reflect exposure to it, not intelligence.
«
97th percentile here (26/30) for age v US white males. (The older you are the more words you’re expected to know.) Fun!
unique link to this extract
Hackers hijack AI accounts and servers to fuel new cyber crime boom • Financial Times
Tom Wilson:
»
Illicit access to AI models and computing power is fast becoming the hottest commodity in the cyber criminal underworld, as hackers seek to harness expensive large language models for extortion, warfare and espionage.
John Hultquist, chief analyst for Google Threat Intelligence Group, said the cyber security unit had observed a major increase in so-called LLM-jacking this year, including the sale of stolen login credentials for public AI tools and the theft of computing resources by groups that want to run their own models for free.
“What we are seeing in the underground is a growing economy associated with access to AI,” Hultquist, a 20-year cyber security veteran, told the FT.
Cheap access to expensive AI gives cyber attackers a financial edge over their targets, who need to use the same AI tools to defend themselves, Hultquist warned.
“The bottom line is all this behaviour gives them a sort of economic or efficiency advantage against us, because they’re going to essentially be able to get their hands on these tokens at a much cheaper rate,” he said.
Marketplaces on the dark web are selling access to AI models from the likes of Anthropic, Google and OpenAI at discounts of up to 97%, Google Threat’s researchers have found. AI subscriptions for the most advanced versions of ChatGPT and Claude can cost as much as $200 per user every month.
Given AI labs monitor for signs of such abuse, some of the sellers even offer “guaranteed access” by committing to provide new credentials at no extra cost if the initial account gets blocked, Hultquist said.
«
Siri AI settlement website now live: Apple to pay some iPhone owners • MacRumors
Joe Rossignol:
»
In May, Apple agreed to pay $250m to settle a US class action lawsuit over Siri AI’s delayed launch, and the settlement website is now live.
Apple will pay an estimated $25 per eligible iPhone to eligible customers, who can submit a claim via the settlement website between September 21 and December 21 of this year. The exact payment may be higher or lower than $25 depending on the total number of valid claims submitted and other factors, with the maximum amount set at $95.
Apple denies all of the allegations made in the lawsuit. Apple agreed to the settlement to avoid the costs and risks of a trial.
In June 2024, Apple previewed new Siri capabilities powered by Apple Intelligence, including understanding of a user’s personal context, on-screen awareness, and deeper per-app controls. For example, Apple showed an iPhone user asking Siri about their mother’s flight and lunch reservation plans based on info from the Mail and Messages apps.
Apple advertised those Siri features in product presentations, on its website, in a TV commercial starring actor Bella Ramsey, and elsewhere.
In March 2025, Apple delayed the launch of the personalized version of Siri, leading to the company being hit with a class action lawsuit alleging false advertising.
«
That is pretty rapid turnaround for a class action suit. The US does these things differently. And we now know the value of AI Siri: it’s $25. (Unfortunately this is only for US users. UK users might see a class representative suit gradually edge through our courts.)
unique link to this extract
Massive UK study finds high-mileage electric cars more durable than petrol • Electrek
Jameson Dow:
»
A study of millions of UK vehicles shows that electric cars stay on the road longer than petrol cars, and the gap widens as mileage increases.
With electric cars being a relatively new technology, many people have been skeptical of what their long-term reliability would look like. Despite EVs having fewer moving parts which suggests less opportunity for failure, fear of the new still keeps people worrying.
But electric cars aren’t really that new anymore. The oldest EVs of the modern era have now been on the road for nearly two decades, so we’ve got some data now. And it turns out EVs look pretty good, according to a new study.
The study was conducted using data from millions of annual British MOT tests, the standard test of roadworthiness for cars on British roads. These tests focus on a car’s suspension, brakes, tires, exhaust, and emissions.
It checked how likely cars were to come back for another MOT test – if they didn’t, then that meant the car was no longer on the road (whether it be for an issue with those five categories above, or another issue).
In the beginning of their operational life, there’s little difference between roadworthiness of petrol and electric vehicles. EVs actually experience slightly more problems in the early years of ownership, but that’s because EVs tend to put on more miles early than petrol cars do.
When comparing on a mileage basis, EVs have very similar failure rates in the early years, differing by less than a percentage point. But that changes once you get to higher mileage numbers – past 40,000 miles lifetime.
At that point, petrol cars are less reliable, and the reliability only gets worse over time. By the time a petrol car reaches 120,000 miles, it’s 47% more likely to have failed than an EV (16% failure rate vs. 23.5% failure rate).
«
40,000 miles doesn’t sound like a lot? Given that most insurers offer 10,000 annually as the normal expectation, that’s only four years. But there’s so much more to go wrong with internal combustion engines that the rising failure rate ought not to be surprising.
unique link to this extract
Germany maps out fossil fuel exit, backing EVs, heat pumps • Reuters
Kirsti Knolle:
»
Germany’s cabinet approved a roadmap on Wednesday to phase out coal, oil and gas by 2045, reaffirming a climate strategy focused on electrification despite a minister’s call for more flexibility on cars, heating and carbon pricing.
An environment ministry spokesperson said the plan would be presented at the United Nations General Assembly in New York later on Wednesday.
Germany would be the third country after France and the Netherlands to adopt a national roadmap under an international effort to move away from fossil fuels, the spokesperson added.
Spiegel magazine earlier on Wednesday reported that, under the plan, Germany will examine whether an agreed deadline to end coal-fired power generation by 2038 can be brought forward to 2035.
In transport, the plan foresees EU fleet emissions rules to lead to battery-electric cars making up 100% of new passenger car registrations by 2035, Spiegel said.
That contrasts with demands from some conservatives, including Economy Minister Katherina Reiche, to preserve a larger role for combustion-engine vehicles using alternative fuels.
«
Incredible how the Green Party in German, through their overreaction to nuclear power, led to coal burning – about as bad for the atmosphere as you can imagine – continuing far past the necessary date. The UK stopped burning coal for power generation at the end of September 2024.
unique link to this extract
The death of apps has already begun • Web Designer Depot
Noah Davis:
»
We spent the last 15 years installing apps for everything. An app for email. An app for music. An app for banking. An app for editing photos. An app for ordering dinner. An app for booking flights. An app for spreadsheets. An app for taking notes. An app for figuring out which other apps we should install.
AI could make much of that feel ridiculous.
Not because AI is going to magically destroy every software company tomorrow, but because the fundamental idea behind the app—the idea that humans should learn an interface designed around a company’s software—suddenly looks surprisingly old-fashioned. The future might not be 200 better apps. It might be one intelligent interface that uses those 200 apps for you.
And the interesting part is that this isn’t entirely science fiction anymore. Apple is making app actions and data available to Apple Intelligence outside the apps themselves through App Intents. OpenAI has moved external apps directly into ChatGPT, while Google’s Gemini can interact with connected apps to perform tasks. The pieces of the post-app world are quietly being assembled while we’re still arguing about whether the Settings button should be in the top-right corner.
Think about what happens when you want to accomplish something today. You decide what you want to do, figure out which app can do it, open that app, navigate its interface, remember where the right feature lives and enter the information in whatever format the software expects. Then, very often, you repeat the whole process in another app.
We’ve become so accustomed to this that we barely notice how strange it is. Want to organize a vacation? You might bounce between a browser, maps, an airline site, a hotel app, your calendar, email, Notes, a weather service and perhaps a spreadsheet. The human is effectively the API connecting everything together.
AI flips that relationship around. Instead of opening eight pieces of software, imagine saying: “Plan a five-day trip to Tokyo in April. Keep the hotel under $250 a night, avoid overnight flights, put everything on my calendar and make me a packing list based on the forecast.”
That’s not really an app interaction anymore. That’s an intention—and intention may become the most important interface in computing.
«
| • Why do social networks drive us a little mad? • Why does angry content seem to dominate what we see? • How much of a role do algorithms play in affecting what we see and do online? • What can we do about it? • Did Facebook have any inkling of what was coming in Myanmar in 2016? Read Social Warming, my latest book, and find answers – and more. |
Errata, corrigenda and ai no corrida: none notified








