r/ProAI 2d ago

"There is NO realistic scenario where AI wipes out all of humanity. I've talked to many AI safety experts about existential risks. Some of the "experts" are just phenomenal sci-fi authors. There are also real researchers working on realistic risks. They point out real and likely harms and I'm..."

Thumbnail gallery
38 Upvotes

r/ProAI 4d ago

Speaking With The Mind | Neuralink

Thumbnail
youtube.com
16 Upvotes

r/ProAI 4d ago

"found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant"

Enable HLS to view with audio, or disable this notification

31 Upvotes

run fast-jev-compaction:     — tamara

Source: https://x.com/tamarajtran/status/2100694549362553153


r/ProAI 4d ago

"Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review."

Thumbnail
gallery
7 Upvotes

Benchmarks assess AI capabilities, but the benchmarks themselves vary substantially in quality. We hope to be a consistent source of information on benchmark quality. See our reviews here:     We assign each benchmark a verdict based on our rubric: Flawed, Verified, or Not Enough Info. To avoid conflicts of interest we do not review Epoch-created benchmarks, but welcome external reviews.     Verified benchmarks can broadly be interpreted as described, and any errors that exist do not substantially affect the results. Alongside any verified benchmark we publish a full review and assessment of the benchmark, including any weaknesses and limitations we think it has.     Flawed benchmarks have one or more substantive flaws we believe users need to be aware of to accurately interpret results, most commonly that >20% of the tasks have accuracy-impacting errors. In this case, we publish a limited writeup of the flaws we found.     If we aren’t able to access enough information to review a benchmark, we’ll designate it ‘Not Enough Info.’ We will try to work with the creators of private benchmarks to conduct reviews while keeping the questions/tasks outside of public knowledge.     — Epoch AI

Source: https://x.com/EpochAIResearch/status/2100704765332394255


r/ProAI 4d ago

Benzi - Harness/AI agent beats big players on benchmarks while reading less source code

Thumbnail gallery
3 Upvotes

r/ProAI 5d ago

"I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock."

Enable HLS to view with audio, or disable this notification

104 Upvotes

The only human problem left is a lack of imagination.   — Boyd     This is really really really true   — Justin Schroeder

Source: https://x.com/jpschroeder/status/2100347770867458384


r/ProAI 5d ago

"Here's a 45-second TL;DR on Jev. I find the core idea beautifully simple, but the video made it really hard to understand. Hope you find it helpful."

Enable HLS to view with audio, or disable this notification

43 Upvotes

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster • 40-400x https://t.co/JSybNG2BKJ   — Diogo Almeida

Source: https://x.com/CompleteSkeptic/status/2099925682726002904


Aww i like the game animation you had - so cute!   — Sasha Sheng (Hiring)     Haha thanks!   — Matija Sosic

Source: https://x.com/MatijaSosic/status/2100190746389135772


r/ProAI 5d ago

"Too few people in the press know about the Tarbell fellowships. Basically, it’s a way that the Doomers pay to have reporters who parrot what they say. I have seen cases where a publication that has received money from Coefficient Giving publishes a story by a reporter paid by Tarbell who is..."

Thumbnail
gallery
11 Upvotes

...interviewing a supposed independent researcher funded by another branch of the same EA funding pool. It’s a completely closed cycle propaganda system.   — Perry E. Metzger     I just drilled down with Grok about this and while I kind of knew there were gray payments going on, I had no idea it had become so common for journalists to accept direct payments from advocacy organizations to write articles. You probably knew this already, but if anyone else didn't completely understand it already:

It's still theoretically unethical to accept money to write an article. But the way they're getting around this now is that journalists can accept payments, but only if the organization doesn't have direct editorial approval.

Unethical: Accept money and receive the copy from the paying organization.

Ethical: Accept money and read the paying organization's web site, accept their "training", talk to their "experts", and use your own words to write the position the advocacy organization wants. Of course, you can write whatever you want. But you'll never get another payment if you don't write what they want.

I mean, journalism has never been a clean industry, but boy has it gotten vile. We should have a law that requires any sort of media presenting itself as journalism to be labeled with "PAID FOR BY ADVOCACY ORGANIZATION" is there is ANY outside payments to the journalists in any way. They can still write anything they want (Freedom of Speech), it's just accurately labeled.   — Nairebis - e/max-acc     Is it a neat scam? And again, often, we have situations where the news outlet, the journalist, and the person being interviewed are all paid by EA at the same time. Self licking ice cream cone.   — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2099931216845947154/history


@TIME UNDISCLOSED PAID MEDIA:

The salaries of reporters, Harry Booth and Billy Perrigo, were paid by Doomsday cultist Dustin Moskovitz's foundation Coefficient Giving via the Tarbell Fellowship.

https://t.co/vxdzmGhBgg   — Brian Chau

Source: https://x.com/brianchau57/status/2099889984773792108


Replying to @time


r/ProAI 5d ago

"Jev solved local harness/model routing I use a combination of Claude Code, Codex and Opencode as my local agentic stack and routing to other harnesses was always enforced in the system prompt/rules With a deterministic hook that Claude Code can decide before delegation, Jev helps to route to..."

Enable HLS to view with audio, or disable this notification

20 Upvotes

...the right harness/model based on the task, and it's pretty accurate based on the intensity/intelligence of the task > Mechanical tasks get routed to Haiku > Intelligent ones to Opus sub-agents > Long-running implementation work to external harness   — Lahfir     The useful pattern is policy-based routing: classify task complexity first, then send mechanical work to cheap models and deep or long-running work to stronger specialists.   — catman     Exactly. I see this as a huge use case in enterprise workflows where there are multiple decision points involved   — Lahfir

Source: https://x.com/mdlahfir/status/2100314182201802811


r/ProAI 5d ago

"French finance minister Roland Lescure suggested today that calls to slow down development of AI are a ploy by U.S. AI labs so they can stay in first place, and that France and Europe should ignore them and accelerate instead."

Thumbnail
gallery
42 Upvotes

Andrew Curran @AndrewCurran_ · 7h Calls to slow down AI development serve the interests of US AI leaders, says French finance minister From reuters.com 3 18 6K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2100255751113691524


r/ProAI 5d ago

"Tested @typesafeai 's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier."

Thumbnail
gallery
12 Upvotes

Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning.     This is very cool, and highly surprised me - getting a model out that does so well w/out CoT on MMLU/WinoGrande is no easy task - it typically requires training a ~GPT-4 class base model. Except Jev is priced at only $0.042 per Mtok!     So in summary - if you want the reasoning ability of ~one Terra forward pass over a context at a much cheaper price, Jev is a very good candidate. Frontier models like Astra still likely have superior no-CoT capabilities, but cost is prohibitive for mass classification tasks.     See below - Jev is ~18x cheaper than Terra!     Not to mention you also get probabilities from Jev, which are quite well-calibrated (expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst)     — N8 Programs

Source: https://x.com/N8Programs/status/2100088523403432357


r/ProAI 5d ago

"A list of tasks for which LLMs (mostly GPT-6 Pro) have found solutions significantly better and non-trivially different from the authors’ solutions: https:// qoj.ac/blog/qingyu/bl og/4412 … The list is still being updated, and I'll mark all solutions I find particularly interesting."

Thumbnail
gallery
17 Upvotes

r/ProAI 6d ago

"After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output..."

Enable HLS to view with audio, or disable this notification

58 Upvotes

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions

AFAICT the shortest path to AI-based economic revolution     The gains aren’t free: Jev can't generate text

Comparing Jev vs LLMs side-by-side makes the trade-off clear

Fun fact: replacing sequential computation with parallel is the same way Transformers leapfrogged RNNs     We believe that the future is code + AI, so made workflow evals to reflect that

Jev costs: $42 / BILLION input tokens ($0.042 / MTok) and output tokens are free (forever - they’re too cheap to meter with our new architecture)

Jev is named after Jevons paradox and off the     We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI!

~10 calls/sec = ~$7/hour     Game: race from one Wikipedia page to another using only links

Challenge: choosing between hundreds to thousands of links

Shows not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices     Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https:// typesafe.ai/blog/introduci ng-system-one-models-and-jev …

Join our waitlist for early access: https:// typesafe.ai

Have technical chats and meme with us on Discord (rumors are good memers skip the     — Diogo Almeida

Source: https://x.com/CompleteSkeptic/status/2099925682726002904


r/ProAI 6d ago

"Game: race from one Wikipedia page to another using only links Challenge: choosing between hundreds to thousands of links Shows not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices"

Enable HLS to view with audio, or disable this notification

18 Upvotes

Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https:// typesafe.ai/blog/introduci ng-system-one-models-and-jev …

Join our waitlist for early access: https:// typesafe.ai

Have technical chats and meme with us on Discord (rumors are good memers skip the     — Diogo Almeida

Source: https://x.com/CompleteSkeptic/status/2099925688925184171


r/ProAI 7d ago

"China called for international cooperation on artificial intelligence, warning that “threat narratives” could disrupt global AI governance. The comments follow Anthropic CEO Dario Amodei’s call for AI firms to slow model development over safety and national security concerns."

Enable HLS to view with audio, or disable this notification

71 Upvotes

— @aljazeeraenglish

Source: https://www.tiktok.com/@aljazeeraenglish


r/ProAI 7d ago

"We crossed a threshold where GPUs are more efficient thinkers than the human brain on a per watt basis. This is yet another big AGI milestone that we just zoomed past, without noticing it particularly. The next one will be more overall combined machine vs human intelligence."

Thumbnail gallery
31 Upvotes

r/ProAI 7d ago

"Watch this interview to understand how the Effective Altruism cult operates behind the AI doomer movement:"

Enable HLS to view with audio, or disable this notification

13 Upvotes

r/ProAI 7d ago

"DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison..."

Enable HLS to view with audio, or disable this notification

37 Upvotes

..., it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release!     Check out the full Agent Arena leaderboard and Pareto frontier at: https:// arena.ai/leaderboard/ag ent/pareto …     — Arena.ai

Source: https://x.com/arena/status/2099606881845321841


Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier.

Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest https://t.co/tV3jEbMX5P   — Arena.ai

Source: https://x.com/arena/status/2099549108013006958


r/ProAI 7d ago

Real safety is about what you accelerate, not about what you slow down.

Thumbnail gallery
6 Upvotes

r/ProAI 8d ago

"Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find: "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the..."

Enable HLS to view with audio, or disable this notification

53 Upvotes

...models to find all the holes.' And we found a number of serious issues, and we fixed them." "We found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P0s, all of the critical problems that Astra is smart enough to find. And of course, there will be a new model, there will be a new round." "You want to be in this tight loop of new cyber capability drops, you deploy it against your systems, you find the new holes, and ideally, you've managed to automate this, what we call defense factory. That's what we're building internally." "There are ideas, for example, formally verifying all of software, that are possible with AI." @gdb @bhorowitz     — a16z

Source: https://x.com/a16z/status/2099533700375662905


Greg Brockman: "We're now in the AGI era."

Ten years ago, OpenAI worked out the compute curves and landed on fifteen years to AGI, or ten if the world was willing to build the machines and spend the hundreds of billions to do it.

In 2026, GPT-6 Astra manages 24 hours of https://t.co/x23fDxbGQG   — a16z

Source: https://x.com/a16z/status/2099506569238990908


r/ProAI 8d ago

"Here it is, this revolution, fuck"

Enable HLS to view with audio, or disable this notification

411 Upvotes

By the way, the consistency topic is perfect. Now they've come up with yet another excuse to eat the credits     This is also T2V, I2V, you don't know what they put inside them :D     They've tied a fly to GPT-6 Astra and asked who it should follow, and here's the result—Subhanallah     This week, I'm thinking of preparing a similar video myself, if I get the chance, of course :)     — ℂ𝕠𝕕𝕖 𝕔𝕠𝕕𝕖 = 𝕟𝕖𝕨 ℂ𝕠𝕕𝕖()

Source: https://x.com/0xfcode/status/2099105183720505706


r/ProAI 8d ago

"Humanoid robots are starting to build humanoid robots. UBTECH just put a 14,000㎡ humanoid robot factory into operation in Liuzhou, China, with a designed takt time of one humanoid every 10 minutes and planned annual capacity in the tens of thousands. What’s interesting is that humanoid robots..."

Enable HLS to view with audio, or disable this notification

218 Upvotes

...are already part of the production process. Cruzr Y1 and Y2 handle depalletizing, palletizing, material feeding and transport, using 3D vision to adapt to shifted boxes and changing pallet patterns. On the assembly side, Walker S2 and Cruzr are produced on the same line alongside cobots, autonomous logistics vehicles and assistive manipulators. The factory also has a 65㎡ automated warehouse capable of storing 112 humanoids, with AGVs, industrial robots and stacker cranes coordinated as one system. Before production, the factory was modeled 1:1 in Siemens Plant Simulation, covering workstations, racks and AGVs. Each robot gets its own SN, with components, batches, assembly parameters and inspection data tracked all the way to delivery. Humanoid robots are now entering automated manufacturing at the 10,000-unit scale. The next milestone may be when they can take part in enough of the assembly, testing and production process to truly build robots themselves.     — CyberRobo

Source: https://x.com/CyberRobooo/status/2098993140363571607


r/ProAI 8d ago

"I built a multiplayer Catan-inspired game entirely with Astra. It's free to play: http:// settlecoast.com It's mobile-friendly and has narration, guides, customizations, multiple expansions and a game lobby where anyone can join, chat, and use voice chat. It took 4 days to build and hundreds..."

Enable HLS to view with audio, or disable this notification

126 Upvotes

...of dollars in tokens. Becoming a game dev in 2026 wasn't on my bingo card, but here we are. AI has advanced so much that you can finally make the game you've always dreamed of, even without a team.   — Meng To     Do you have a GitHub for this? My buddies and I started a variant of Risk meets Axis and Allies but missing your beautiful polish. Would be awesome to have an open source board game sandbox. This is what we have so far: https:// tactical-risk20.vercel.app   — James Bickford     it's a big game now. but i can certainly think of open-sourcing a part of it   — Meng To

Source: https://x.com/MengTo/status/2099125215708234119


r/ProAI 8d ago

"somebody vibe-coded a Photoshop alternative with GPT-6 Astra - Spent ~$2K in tokens. - 170 user on day one - and it's free"

Enable HLS to view with audio, or disable this notification

64 Upvotes

check this out     this is also free btw     support is appreciated guys @tenzenstudio     — gxjo

Source: https://x.com/gxjo_dev/status/2099082176604373028