Start Up No.2744: can an AI beat Pangram’s detector?, Siri AI put to the test, AI-produced sitcom airs, charge up!, and more


Lifting lots of weights might reduce your “health age” as measured by various systems, but does that have any meaning? CC-licensed photo by Usodesita on Flickr.

You can sign up to receive each day’s Start Up post by email. You’ll need to click a confirmation link, so no spam.


A selection of 9 links for you. What about wealthy or wise? I’m @charlesarthur on Twitter. On Threads: charles_arthur. On Mastodon: https://newsie.social/@charlesarthur. On Bluesky: @charlesarthur.bsky.social. Observations and links welcome.


Can I get an AI to beat Pangram? • Ad Rem

Remi:

»

I started off by asking GPT if it thought it was among the most human sounding AIs. It said yes.

So, I chose Claude.

And after I chose Claude, I checked Google and other AI to see if it could step up to the challenge and surprisingly, the other AIs glazed Claude as an AI among AIs— Astonished by how human it sounds. But we are the ones who are going to put that to the test.

I am going to be using the Sonnet 5 model on high effort as well.

Now that that’s out of the way, I am mostly going to give the AI some general and some specific topics to write about to see if Pangram flags it, then, I will have the AI modifies what it says continuously in attempt to get around Pangram’s flagging. If Pangram gets it right, it should flag the text as 100% AI (or near that) almost every single time. Now, Pangram has lower confidence rates when it comes to shorter text, so I will give Claude a baseline of at least 400 words to write

If you are interested in how Pangram works, read this article and further, read this article on why AI text is, as Pangram claims, distinctive from human writing.

So, this game will be one about what Pangram calls “Adversarial prompting”

Both of these articles are important because I will be using them to inform Claude on how to avoid Pangram’s detection, and if the team of Claude and I are really successful, then Claude will get a 100% human written score from Pangram or, after some amount of attempts it will get a more human score, easy enough right? Let’s do this.

«

He gets Claude to write a short story: Pangram nails it as 100% AI. He gets it to tweak it: zapped again. And when he tells Claude to try to get smarter, Claude goes all HAL 9000 on him, regretting that he can’t edit the pod bay doors, Hal, no matter how you phrase it.

Pangram turns out to be really good at identifying AI-written content, based on this.
unique link to this extract


125 tests of the new AI Siri • Pogue’s Posts

David Pogue really put it through its paces:

»

You can say things like, “What did my wife say I was supposed to do with the plants?” Or, “What was the show that Casey said I should watch?” Or, “What’s my flight number to Austin?” And it pops right up.

You may not even care where that information actually lives—some text message, some email, some note. All you know is that you didn’t have to go scrolling your messages back to A.D. 700 to find the answer.

The hardest part will be getting out of the habit of opening apps, and remembering that you have Siri. Give it a few tries. Learn to trust it. You’ll save incredible amounts of time and fumbling.

I used the beta-test version all summer, growing more and more excited. Now that the final version is out, I wanted to see how often it gets things right. It is AI, after all. So I ran it through 125 tests.

Some of these tasks are things Apple said are possible. Some ideas, I picked up from people on Reddit trying stuff. Some, I just wondered if Siri could do.

It failed a handful of times. But a few other times, it absolutely blew me away. It did amazing things I bet you never suspected Siri could do. Note, for example, items 24, 25, 28, 30, 61, 100, and 124.

(Note: To get started, open Settings → Siri and turn on Siri AI. For now, you may be told you’re on the waiting list; Apple can’t unleash it to 2.5 billion devices all at once.)

OK, here we go.

«

I think strictly it’s called “Siri AI”, but anyway. What we haven’t seen yet (does it exist?) is a side-by-side with an Android phone running Gemini.
unique link to this extract


Trump administration moves to gut a core endangered species protection • The New York Times

Catrin Einhorn:

»

In what would be a seismic shift, the Trump administration is quietly moving to strip a core protection of the Endangered Species Act, according to an internal memo obtained by The New York Times.

Under a new interpretation of the law’s language outlined in the memo, the accidental killing or injury of an animal would no longer be considered illegal; only actions specifically intending to target an animal would.

The change would be transformative because, in the United States, the killing of endangered animals almost always happens incidentally, in the course of economic activity. Until now, industries and individuals could be held accountable for those foreseeable deaths and, as a result, they have been incentivized to avoid them.

At the center of the change is what it means to “take” an endangered animal, something that has long been prohibited under the act.

“A vessel that inadvertently strikes a whale has not taken it, because the vessel’s course was not set against the whale,” states the memo, which was dated Sept. 14 and was distributed to U.S. Fish and Wildlife Service field offices. “Felling a tree is not a take of the bats roosting in it unless the tree is felled for the purpose of killing or capturing them.”

Currently, enforcement of the Endangered Species Act relies on a permitting system that seeks to balance protecting imperiled animals with allowing business to operate.

For example, if developers wanted to build houses on a tract of land that is home to endangered salamanders, they would have to apply for a permit. In exchange for implementing some measures to protect the salamanders, and perhaps paying for salamander conservation elsewhere, the Fish and Wildlife Service would issue what’s known as an incidental take permit, absolving them of killing a certain number of salamanders.

The memo, which was signed by Brian Nesvik, the director of the Fish and Wildlife Service, appears to upend that entire system.

«

There’s just no element of life that the Trump administration isn’t making worse, either accidentally or intentionally.
unique link to this extract


Your “health age” is fake • The Verge

Victoria Song:

»

It’s true that through a healthy diet, good sleeping habits, and regular exercise, your physiological condition might be better than that of other people who are the same age and fit within a similar demographic. Depending on how a company defines its various longevity-related age scores, your results can identify areas where you can improve. For example, my Whoop age is somewhat cursed because it heavily weights strength training workouts in the algorithm. The last time I tested Whoop, I wasn’t going super hard on lifting heavy and perhaps was a bit too into cardio. If I wanted my Whoop age to improve, lifting weights would’ve been an easy way to tilt that number in my favor.

Objectively, recalibrating my training habits to be more balanced would be a good thing for my long-term health. If you leave it exactly at that, there’s nothing wrong with this type of feature. But culturally speaking, that’s not how these “health age” scores work.

The way I’ve seen people talk about this metric largely hews toward fluffing up your ego. If your health age is older, you probably don’t talk about it. It’s “shameful” even though, really, it’s not. That’s because culturally — at least in the West — getting old isn’t considered a good thing. If your health age is younger, it becomes something to bandy about.

Apple’s vice president of health, Sumbul Desai, even says in this Instagram Reel that Health Age could be something you brag about to your friends. That explains why tech entrepreneur Bryan Johnson, famous for his quest to never die, would post this absurd take on X about how aliens would struggle to distinguish him from his teenage son, alongside a chart titled “Is Bryan Johnson More 18 or 49?”

Johnson is 49; he looks a little younger, maybe, and his metrics might be comparable to someone much younger on average, but his current state is ultimately just what 49 looks like for him.

Furthermore, your wearable-determined age is oftentimes not capturing the whole picture. You could game the Whoop Age metric to the letter, get your annual physical, and find out that your cholesterol level is through the roof. A reader once related this exact story to me. You could optimize all the controllable lifestyle health factors — diet, exercise, and sleep — like Bryan Johnson and still end up with autoimmune gastritis, as he has. Some of your health is encoded in your genetics, and while you can better the odds, your ultimate fate is not always something you can decide.

«

“Health age” really is trying to gamify a game you cannot win.
unique link to this extract


Sitcom produced entirely by AI is released on YouTube • The Irish News

Sean O’Grady:

»

A sitcom produced entirely by artificial intelligence has premiered on YouTube.

No Big Deal, written by Andrew Dickinson, was produced by UK AI studio ModeLabs.ai.

Set inside the chaotic world of venture capital, No Big Deal follows an eccentric team at investment firm Janus as they search for the next billion-pound business idea – usually with disastrous results.

Dickinson said he originally planned to make the pilot through “traditional television production” but found the process would be “slow and prohibitively expensive”. He added: “When I met the team at ModeLabs, I realised AI offered an entirely new way to bring the project to life.

“Making the pilot wasn’t simply a case of pressing a button. Every shot, camera movement, performance, line reading, location and visual detail had to be carefully planned, refined and recreated through an iterative creative process between the writing and production teams.

“Even while we were making it, the script was rewritten more than 65 times. I’m incredibly proud of what we’ve created together. I still laugh at these characters, and I hope audiences will too.”

Dan Hawthorne, from ModeLabs.ai, said: “Nobody questions whether the CGI in a Marvel film is ‘cheating’. People judge entertainment by how it makes them feel, not by the tools used to create it.

“No Big Deal is brilliantly written and genuinely funny. AI isn’t the punchline – it’s simply the production technology that made an ambitious independent comedy possible.

«

So they made a version of Silicon Valley, located in London, but far, far, far worse. The characterisation is poor, the plots are bad, the scenes don’t make sense. There’s a half-hour clip embedded on the page with this story which demonstrates that the key to comedy is writing, and the secret is editing, editing, editing. Rewriting the script 65 times is about one-fifth as many as required. (It also deserved the headline for the AI-generated “Odyssey”: half-hour AI sitcom is half an hour too long.)

Watch five minutes of Silicon Valley to wash your brain out. (Try this 90-second clip, chosen at random, and note its effortless moments of physical comedy, besides the dialogue.)
unique link to this extract


AI existential risk probabilities are too unreliable to inform policy • AI as Normal Technology

Arvind Narayanan and Sayash Kapoor, writing in July 2024:

»

In this first essay, we look at one type of evidence: probability estimates. The AI safety community relies heavily on forecasting the probability of human extinction due to AI (in a given timeframe) in order to inform decision making and policy. An estimate of 10% over a few decades, for example, would obviously be high enough for the issue to be a top priority for society. 

Our central claim is that AI x-risk forecasts are far too unreliable to be useful for policy, and in fact highly misleading.

If the two of us predicted an 80% probability of aliens landing on earth in the next ten years, would you take this possibility seriously? Of course not. You would ask to see our evidence. As obvious as this may seem, it seems to have been forgotten in the AI x-risk debate that probabilities carry no authority by themselves. Probabilities are usually derived from some grounded method, so we have a strong cognitive bias to view quantified risk estimates as more valid than qualitative ones. But it is possible for probabilities to be nothing more than guesses. Keep this in mind throughout this essay (and more broadly in the AI x-risk debate).

If we predicted odds for the Kentucky Derby, we don’t have to give you a reason — you can take it or leave it. But if a policymaker takes actions based on probabilities put forth by a forecaster, they had better be able to explain those probabilities to the public (and that explanation must in turn come from the forecaster). Justification is essential to legitimacy of government and the exercise of power. A core principle of liberal democracy is that the state should not limit people’s freedom based on controversial beliefs that reasonable people can reject. 

Explanation is especially important when the policies being considered are costly, and even more so when those costs are unevenly distributed among stakeholders. A good example is restricting open releases of AI models. Can governments convince people and companies who stand to benefit from open models that they should make this sacrifice because of a speculative future risk?

«

Nothing has changed, except we’re taking about aliens landing a lot more.
unique link to this extract


AI eyes in the sky: new satellites and artificial intelligence are transforming wildfire detection • The Guardian

Ben Tracy:

»

The fires in Spokane, Washington, this summer are the latest reminder that when a wildfire is detected can be just as important as where it starts. A blaze found within minutes can often be contained before it spreads. Wait too long, and a small spark can become a major disaster.

A new effort to shorten that timeline began in July when a SpaceX rocket carried the first three FireSat satellites into orbit. The trio is the first part of a planned network of 50 satellites built solely to spot wildfires while they are still small enough to stop.

FireSat is designed to fill an important gap. Existing weather and Earth-observing satellites can detect large fires, but they often lack the resolution or revisit the same location too infrequently to spot small fires shortly after they ignite. FireSat is designed to detect fires as small as a beach bonfire and, once the full constellation is deployed, scan every point on Earth about every 20 minutes.

The satellites use infrared sensors and artificial intelligence to search for heat that could signal a new wildfire. The system compares new images with earlier ones and accounts for factors such as weather conditions and nearby sources of heat before sending alerts to emergency responders. The goal is to identify real wildfires while reducing false alarms.

FireSat is designed to work alongside a growing network of cameras on the ground that use artificial intelligence to search for smoke and heat. Together, the systems could give firefighters their earliest warning yet that a fire has started.

…Pano AI has installed more than 1,400 cameras across 17 states. The cameras continuously scan the landscape from towers and mountain tops. Artificial intelligence reviews every image, searching for smoke during the day and heat after dark.

“The AI algorithms are being run continuously on that imagery to say ‘smoke, not smoke,’” Satyam said. “At night-time, we’re looking at heat signatures, so heat versus cold. And then we’re able to zoom in and validate that it is a potential fire start.”

«

unique link to this extract


I sent Google proof of a bot farm. They called it “normal user behaviour” • Dayzle

Nick Abe:

»

Last week I wrote about spending $220 on Google app ads and finding out 60% of the installs were robots. I sent Google the evidence and promised to report back on the refund. Google has answered, and to understand why I feel like I’m going crazy, you probably need the other half of my Google Ad(Mob) experience.

We show ads in Dayzle through Google AdMob. A few days after our soft launch, Google suspended the account for “self-clicking”: a month with no ads, no explanation of which clicks, and no appeal.

So I went looking. Our entire AdMob history was 255 ads shown and 15 clicks, every one of them on an iPhone or iPad in Canada. It turns out that on certain iOS devices, Google’s full-screen ads show no close button, whatever the ad is. If you can’t close an ad, you tap the screen, and tapping the screen clicks the ad. On one friend’s iPhone 12, that happened with every full-screen ad they saw. On another friend’s iPad, the ads never showed a close button either, and sometimes froze the whole device, so the only way out was to restart it. I only know about those two because I know them personally.

That means at least two devices had a 100% click-through rate: every ad they were shown got clicked. I understand why that looks suspicious. It looks exactly like me sitting in my basement, clicking my own ads to make a few pennies. By my count, somewhere between 5 and 15 of our 15 clicks came from this bug.

Total earnings before the suspension, from real ads shown to real people: $3.77. Canadian, so less than $2.75 in freedom dollars.

I contacted Google’s ad support, and they took it seriously. They confirmed the missing close button was a known issue and gave it an internal bug number. Then they asked me to help them debug it, by recording the app’s network traffic while the bug happens.

Google, a company with a market cap of about $4.2 trillion at the time of writing and about 200,000 employees, needed my help. One of the largest companies in the world, with a bug in one of its biggest ad platforms, asked a guy who makes puzzle games to help track it down. What reality are we living in? Is this the darkest timeline?

«

unique link to this extract


More than 3,000 new EV charge points planned across UK • The Herald

Rose Moncur:

»

More than 3,000 new electric vehicle charge points are set to be installed across the UK and Ireland as part of a major expansion of public charging infrastructure.

The project is being led by Ireland-headquartered EZO and financed through a £150 million senior debt facility backed by lenders including Aberdeen Investments and Standard Life.

It includes a 20-year concession covering four counties in northern Scotland and will deliver more than 100,000kW of charging capacity across the UK and Ireland.

The funding will support the design, installation, operation and maintenance of the charge points.

The infrastructure is expected to generate more than £1 billion in electricity sales over the 20-year concession period.

A share of the revenue will go to the public sector, while ownership of the chargers will transfer back to public authorities when the concession ends.

EZO has expanded rapidly in the UK and expects the market to become its primary focus by 2027.

«

The money generated is interesting: £1bn over 20 years is an average of £50m per year from those 3,000 chargers, or about £16,667 per year per charger. This link suggests that’s rather more than the average, which ranges from £8,000 to £15,000 on typical usage.
unique link to this extract


• Why do social networks drive us a little mad?
• Why does angry content seem to dominate what we see?
• How much of a role do algorithms play in affecting what we see and do online?
• What can we do about it?
• Did Facebook have any inkling of what was coming in Myanmar in 2016?

Read Social Warming, my latest book, and find answers – and more.


Errata, corrigenda and ai no corrida: none notified

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.