

I’m okay with the downvotes. There’s a stronger anti-AI sentiment now than even just a month ago, and on the whole I think that’s going to be a good thing.


I’m okay with the downvotes. There’s a stronger anti-AI sentiment now than even just a month ago, and on the whole I think that’s going to be a good thing.


All fair questions. I can imagine OpenAI would cheat in any way they can, but I suppose I’m putting some trust in ARC Prize, or specifically Francois Chollet. I’d be happy to be wrong about this because the implications of it being real are worrying (and I try not to add to the hype/doom). Time will tell I suppose, but I still think it’s worth sharing compared to other benchmarks.


Yes, I’m aware many benchmarks are easily gamed and rely mostly on knowledge acquisition rather than true reasoning. I wouldn’t post other benchmark results either.
However, ARC-AGI-3 is unique in that it tests systems using dynamic, original game environments where the AI has to learn each game from scratch. Importantly, the score is based on action efficiency, not simply beating each game.
So this benchmark is a bigger signal than most. It should still be judged with skepticism, certainly and doesn’t necessarily indicate a fully general intelligence, but it seems like a step-change to me.
I’d encourage you to read more about the benchmark, including the technical paper, if you’re interested.


Yes, it was released in March. One important point is that these results are based on the semi-private test set. Perhaps it’s best to wait until it is measured against the fully private test set, but that might not happen for months.


I agree with most of this. I’ve also said in the past that LLMs cannot think, and I think that’s still true for most models. The reason ARC-AGI-3 is interesting is that it was specifically designed to test reasoning, adaptability, novel problem solving, planning, memory, etc. So it was a surprise to me that Astra was able to defeat it so effectively, and that Astra invents algebras for each novel task.
But I agree we can’t trust OpenAI if these results are self-reported, and we may not be able to trust the ARC Prize Foundation fully either. Extraordinary claims require extraordinary evidence, so we need replication, transparency, and proper open science to confirm things.
I also agree with ARC Prize’s conclusion, that there are still capabilities any AI system would need to demonstrate before we can claim a full general intelligence.


I guess this is being down-voted just because it feels pro-AI. Believe me I get it, you can look at my anti-AI post history. But I think it’s important to also keep an eye out for real signals towards general intelligence, which is why I wanted to share this. I’m not celebrating it, just bringing awareness.


I guess this is being down-voted just because it feels pro-AI. Believe me I get it, you can look at my anti-AI post history. But I think it’s important to also keep an eye out for real signals towards general intelligence, which is why I wanted to share this. I’m not celebrating it, just bringing awareness.


Prediction is very hard, but if we’re talking about centuries, I could see it happen.
There was a growing vegan trend a decade or so ago, but I think that’s subsided largely because of political right-wing, conservative trends but also because people care less about ethical issues when they are struggling with basic needs and global political turmoil. If the future leans back towards the left, veganism might trend back.
Climate change might be a big factor in moving away from industrialized animal farming. it’s just so inefficient. I also feel like it’s becoming more popular to reduce meat consumption at least, so I could see people cutting down significantly while not adopting a vegan or vegetarian label. Dairy is a tough one though, it’s much more ingrained and normalized.
People forget that we already have examples of societies like India that are largely vegetarian. If we already have a success case, I don’t see why that can’t be replicated globally. India fell into it through a cultural and religious accident, but I think the mindset can spread just as well in other countries, so it’s important and worthwhile to keep changing peoples’ minds and normalize plant-based lives.


They basically pass the buck to the individual developer without taking any responsibility themselves.
Debian acknowledges that the legal status of material produced by generative AI systems remains the subject of ongoing discussion in many jurisdictions, including questions relating to copyright, authorship, licensing, and potential reproduction of training material.
The responsibility for every contribution rests with the contributor who submits it, who remains accountable for its technical quality, legal acceptability, and suitability for inclusion in Debian.
“It may be illegal or against FOSS, but that’s up to you to decide, good luck I guess”


Here’s the full petition page if you want to see all the detail: https://www.ourcommons.ca/petitions/en/Petition/Details?Petition=e-7531


The world’s pretty absurd. I’d like avoid participating in the absurdity as best as I can.
It’s pretty absurd that we impregnate cows just to milk them.
It’s pretty absurd that we slaughter pigs who have intelligence on par with three-year-old children.
It’s pretty absurd that 80% of agricultural land is used for livestock even though 83% of our calories come from plant-based foods.
It’s pretty absurd that our cattle herds are so large that they produce multiple gigatonnes of greenhouse gases from their digestion alone.
It’s pretty absurd that we’ve bred chickens to lay 29 times more eggs per year than their wild counterparts.
It’s pretty absurd that we do all of this at subsidized, unsustainable, industrial levels.
It’s pretty absurd that we do this mostly for pleasure, taste, and tradition, not necessity. Especially when the vast majority of us have access to perfectly viable and often delicious alternatives.


Not trying to pick a fight with you, friend. I don’t think billionaires should exist. Just saying it’s a bit more complicated than “billionaires bad”.


If you’re referring to American billionaires, I’d probably agree. But I think this section from the article is more accurate and less to do with billionaires:
Confusion and fear are possible reasons why people are becoming less accepting of gender diversity, said Wayne Bernakevitch, founder of the Regina Civic Awareness Action Network, an organization that lobbied in favour of the Parents’ Bill of Rights.
“They’re concerned with finding a house to live in, you know, how do you pay for the next week’s groceries? And so getting into all of this is just a great distraction in their life,” Bernakevitch said in an interview with CBC’s Blue Sky.
Politicians play a critical role in the creation of these attitudes, Leah Hamilton, a professor of business, communication studies and aviation at Mount Royal University, told Blue Sky.
“We have seen over and over throughout history, especially during periods of perceived or real competition for resources like jobs and housing and other things, that politicians will often create scapegoats,” Hamilton said.


my only regret is i’ve already signed this 🍉


The petition: https://www.ourcommons.ca/petitions/en/Petition/Details?Petition=e-7531
Now at 35,785 signatures.


The petition: https://www.ourcommons.ca/petitions/en/Petition/Details?Petition=e-7531
Now at 35,782 signatures.


ARC Prize maintains multiple tracks around their benchmarks. They have “verified” leaderboards, “community” leaderboards, and they also run the ARC Prize competition.
They update the “verified” leaderboards when they test raw LLMs without sophisticated harnesses. They seem to update this sporadically and only occasionally do press releases or blog posts about new scores. For example, the latest score from Claude Opus 5 (High) is 30% at $20,000, and they didn’t post about that as far as I know. Again, this just the raw LLM without an agentic or world-model harness.
The ARC Prize competition has a harder set of criteria. Participants have to use smaller, open models with a limited compute budget, with open source code, and of course the solutions are verified by ARC Prize at the end of the competition.
The “community” leaderboards, which is what this post is about, are self-reported and not verified by ARC Prize. There are no restrictions on what model is used or limitations on compute. So naturally they aren’t going to make official news releases about those, unless they decide to verify them at some point.
The only reason I chose to post this is that the top solutions seem legitimate, with source code released, and two of them have associated papers.


It think it’s still unwise to talk about these topics in broad terms like AGI and even “intelligence”. We still have to pick the capabilities apart to have useful discussions about them. I agree these games are better tests than many benchmarks, but it’s also important to note that these solutions use a combination of well-designed deterministic harnesses, as well as LLMs. So it’s inaccurate to say that “LLMs have achieved AGI” (not sure if that’s what you were getting at). This feels like an important milestone, but we’ll have to continue to probe for failure cases in other categories of problems.
Aside from emotional intelligence, experience, embodiment, etc., these ARC-AGI-3 solutions all rely on the sandbox being a safe environment to fail. The solutions iterate through the problem thousands of times before coming to a final solution. Many real-world human problems cannot be re-tried safely or efficiently.


Thanks for sharing this!
A link to a login wall is just spam.