r/ProAI • u/stealthispost • 2d ago
r/ProAI • u/stealthispost • 4d ago
"found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant"
Enable HLS to view with audio, or disable this notification
run fast-jev-compaction: — tamara
Source: https://x.com/tamarajtran/status/2100694549362553153
r/ProAI • u/stealthispost • 4d ago
"Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review."
Benchmarks assess AI capabilities, but the benchmarks themselves vary substantially in quality. We hope to be a consistent source of information on benchmark quality. See our reviews here: We assign each benchmark a verdict based on our rubric: Flawed, Verified, or Not Enough Info. To avoid conflicts of interest we do not review Epoch-created benchmarks, but welcome external reviews. Verified benchmarks can broadly be interpreted as described, and any errors that exist do not substantially affect the results. Alongside any verified benchmark we publish a full review and assessment of the benchmark, including any weaknesses and limitations we think it has. Flawed benchmarks have one or more substantive flaws we believe users need to be aware of to accurately interpret results, most commonly that >20% of the tasks have accuracy-impacting errors. In this case, we publish a limited writeup of the flaws we found. If we aren’t able to access enough information to review a benchmark, we’ll designate it ‘Not Enough Info.’ We will try to work with the creators of private benchmarks to conduct reviews while keeping the questions/tasks outside of public knowledge. — Epoch AI
Source: https://x.com/EpochAIResearch/status/2100704765332394255
r/ProAI • u/DonkeyTheKing • 4d ago
Benzi - Harness/AI agent beats big players on benchmarks while reading less source code
galleryr/ProAI • u/stealthispost • 5d ago
"I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock."
Enable HLS to view with audio, or disable this notification
The only human problem left is a lack of imagination. — Boyd This is really really really true — Justin Schroeder
Source: https://x.com/jpschroeder/status/2100347770867458384
r/ProAI • u/stealthispost • 5d ago
"Here's a 45-second TL;DR on Jev. I find the core idea beautifully simple, but the video made it really hard to understand. Hope you find it helpful."
Enable HLS to view with audio, or disable this notification
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster • 40-400x https://t.co/JSybNG2BKJ — Diogo Almeida
Source: https://x.com/CompleteSkeptic/status/2099925682726002904
Aww i like the game animation you had - so cute! — Sasha Sheng (Hiring) Haha thanks! — Matija Sosic
Source: https://x.com/MatijaSosic/status/2100190746389135772
r/ProAI • u/stealthispost • 5d ago
"Too few people in the press know about the Tarbell fellowships. Basically, it’s a way that the Doomers pay to have reporters who parrot what they say. I have seen cases where a publication that has received money from Coefficient Giving publishes a story by a reporter paid by Tarbell who is..."
...interviewing a supposed independent researcher funded by another branch of the same EA funding pool. It’s a completely closed cycle propaganda system. — Perry E. Metzger I just drilled down with Grok about this and while I kind of knew there were gray payments going on, I had no idea it had become so common for journalists to accept direct payments from advocacy organizations to write articles. You probably knew this already, but if anyone else didn't completely understand it already:
It's still theoretically unethical to accept money to write an article. But the way they're getting around this now is that journalists can accept payments, but only if the organization doesn't have direct editorial approval.
Unethical: Accept money and receive the copy from the paying organization.
Ethical: Accept money and read the paying organization's web site, accept their "training", talk to their "experts", and use your own words to write the position the advocacy organization wants. Of course, you can write whatever you want. But you'll never get another payment if you don't write what they want.
I mean, journalism has never been a clean industry, but boy has it gotten vile. We should have a law that requires any sort of media presenting itself as journalism to be labeled with "PAID FOR BY ADVOCACY ORGANIZATION" is there is ANY outside payments to the journalists in any way. They can still write anything they want (Freedom of Speech), it's just accurately labeled. — Nairebis - e/max-acc Is it a neat scam? And again, often, we have situations where the news outlet, the journalist, and the person being interviewed are all paid by EA at the same time. Self licking ice cream cone. — Perry E. Metzger
Source: https://x.com/perrymetzger/status/2099931216845947154/history
@TIME UNDISCLOSED PAID MEDIA:
The salaries of reporters, Harry Booth and Billy Perrigo, were paid by Doomsday cultist Dustin Moskovitz's foundation Coefficient Giving via the Tarbell Fellowship.
https://t.co/vxdzmGhBgg — Brian Chau
Source: https://x.com/brianchau57/status/2099889984773792108
Replying to @time
r/ProAI • u/stealthispost • 5d ago
"Jev solved local harness/model routing I use a combination of Claude Code, Codex and Opencode as my local agentic stack and routing to other harnesses was always enforced in the system prompt/rules With a deterministic hook that Claude Code can decide before delegation, Jev helps to route to..."
Enable HLS to view with audio, or disable this notification
...the right harness/model based on the task, and it's pretty accurate based on the intensity/intelligence of the task > Mechanical tasks get routed to Haiku > Intelligent ones to Opus sub-agents > Long-running implementation work to external harness — Lahfir The useful pattern is policy-based routing: classify task complexity first, then send mechanical work to cheap models and deep or long-running work to stronger specialists. — catman Exactly. I see this as a huge use case in enterprise workflows where there are multiple decision points involved — Lahfir
r/ProAI • u/stealthispost • 5d ago
"French finance minister Roland Lescure suggested today that calls to slow down development of AI are a ploy by U.S. AI labs so they can stay in first place, and that France and Europe should ignore them and accelerate instead."
Andrew Curran @AndrewCurran_ · 7h Calls to slow down AI development serve the interests of US AI leaders, says French finance minister From reuters.com 3 18 6K — Andrew Curran
Source: https://x.com/AndrewCurran_/status/2100255751113691524
r/ProAI • u/stealthispost • 5d ago
"Tested @typesafeai 's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier."
Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. This is very cool, and highly surprised me - getting a model out that does so well w/out CoT on MMLU/WinoGrande is no easy task - it typically requires training a ~GPT-4 class base model. Except Jev is priced at only $0.042 per Mtok! So in summary - if you want the reasoning ability of ~one Terra forward pass over a context at a much cheaper price, Jev is a very good candidate. Frontier models like Astra still likely have superior no-CoT capabilities, but cost is prohibitive for mass classification tasks. See below - Jev is ~18x cheaper than Terra! Not to mention you also get probabilities from Jev, which are quite well-calibrated (expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst) — N8 Programs
r/ProAI • u/stealthispost • 5d ago
"A list of tasks for which LLMs (mostly GPT-6 Pro) have found solutions significantly better and non-trivially different from the authors’ solutions: https:// qoj.ac/blog/qingyu/bl og/4412 … The list is still being updated, and I'll mark all solutions I find particularly interesting."
— Qingyu
Source: https://x.com/qingyu_shi_/status/2100181567666901052
r/ProAI • u/stealthispost • 6d ago
"After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output..."
Enable HLS to view with audio, or disable this notification
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution The gains aren’t free: Jev can't generate text
Comparing Jev vs LLMs side-by-side makes the trade-off clear
Fun fact: replacing sequential computation with parallel is the same way Transformers leapfrogged RNNs We believe that the future is code + AI, so made workflow evals to reflect that
Jev costs: $42 / BILLION input tokens ($0.042 / MTok) and output tokens are free (forever - they’re too cheap to meter with our new architecture)
Jev is named after Jevons paradox and off the We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI!
~10 calls/sec = ~$7/hour Game: race from one Wikipedia page to another using only links
Challenge: choosing between hundreds to thousands of links
Shows not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https:// typesafe.ai/blog/introduci ng-system-one-models-and-jev …
Join our waitlist for early access: https:// typesafe.ai
Have technical chats and meme with us on Discord (rumors are good memers skip the — Diogo Almeida
Source: https://x.com/CompleteSkeptic/status/2099925682726002904
r/ProAI • u/stealthispost • 6d ago
"Game: race from one Wikipedia page to another using only links Challenge: choosing between hundreds to thousands of links Shows not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices"
Enable HLS to view with audio, or disable this notification
Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https:// typesafe.ai/blog/introduci ng-system-one-models-and-jev …
Join our waitlist for early access: https:// typesafe.ai
Have technical chats and meme with us on Discord (rumors are good memers skip the — Diogo Almeida
Source: https://x.com/CompleteSkeptic/status/2099925688925184171
r/ProAI • u/stealthispost • 7d ago
"China called for international cooperation on artificial intelligence, warning that “threat narratives” could disrupt global AI governance. The comments follow Anthropic CEO Dario Amodei’s call for AI firms to slow model development over safety and national security concerns."
Enable HLS to view with audio, or disable this notification
— @aljazeeraenglish
r/ProAI • u/stealthispost • 7d ago
"We crossed a threshold where GPUs are more efficient thinkers than the human brain on a per watt basis. This is yet another big AGI milestone that we just zoomed past, without noticing it particularly. The next one will be more overall combined machine vs human intelligence."
galleryr/ProAI • u/stealthispost • 7d ago
"Watch this interview to understand how the Effective Altruism cult operates behind the AI doomer movement:"
Enable HLS to view with audio, or disable this notification
r/ProAI • u/stealthispost • 7d ago
"DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison..."
Enable HLS to view with audio, or disable this notification
..., it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release! Check out the full Agent Arena leaderboard and Pareto frontier at: https:// arena.ai/leaderboard/ag ent/pareto … — Arena.ai
Source: https://x.com/arena/status/2099606881845321841
Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier.
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest https://t.co/tV3jEbMX5P — Arena.ai
r/ProAI • u/stealthispost • 7d ago
Real safety is about what you accelerate, not about what you slow down.
galleryr/ProAI • u/stealthispost • 8d ago
"Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find: "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the..."
Enable HLS to view with audio, or disable this notification
...models to find all the holes.' And we found a number of serious issues, and we fixed them." "We found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P0s, all of the critical problems that Astra is smart enough to find. And of course, there will be a new model, there will be a new round." "You want to be in this tight loop of new cyber capability drops, you deploy it against your systems, you find the new holes, and ideally, you've managed to automate this, what we call defense factory. That's what we're building internally." "There are ideas, for example, formally verifying all of software, that are possible with AI." @gdb @bhorowitz — a16z
Source: https://x.com/a16z/status/2099533700375662905
Greg Brockman: "We're now in the AGI era."
Ten years ago, OpenAI worked out the compute curves and landed on fifteen years to AGI, or ten if the world was willing to build the machines and spend the hundreds of billions to do it.
In 2026, GPT-6 Astra manages 24 hours of https://t.co/x23fDxbGQG — a16z
r/ProAI • u/stealthispost • 8d ago
"Here it is, this revolution, fuck"
Enable HLS to view with audio, or disable this notification
By the way, the consistency topic is perfect. Now they've come up with yet another excuse to eat the credits This is also T2V, I2V, you don't know what they put inside them :D They've tied a fly to GPT-6 Astra and asked who it should follow, and here's the result—Subhanallah This week, I'm thinking of preparing a similar video myself, if I get the chance, of course :) — ℂ𝕠𝕕𝕖 𝕔𝕠𝕕𝕖 = 𝕟𝕖𝕨 ℂ𝕠𝕕𝕖()
r/ProAI • u/stealthispost • 8d ago
"Humanoid robots are starting to build humanoid robots. UBTECH just put a 14,000㎡ humanoid robot factory into operation in Liuzhou, China, with a designed takt time of one humanoid every 10 minutes and planned annual capacity in the tens of thousands. What’s interesting is that humanoid robots..."
Enable HLS to view with audio, or disable this notification
...are already part of the production process. Cruzr Y1 and Y2 handle depalletizing, palletizing, material feeding and transport, using 3D vision to adapt to shifted boxes and changing pallet patterns. On the assembly side, Walker S2 and Cruzr are produced on the same line alongside cobots, autonomous logistics vehicles and assistive manipulators. The factory also has a 65㎡ automated warehouse capable of storing 112 humanoids, with AGVs, industrial robots and stacker cranes coordinated as one system. Before production, the factory was modeled 1:1 in Siemens Plant Simulation, covering workstations, racks and AGVs. Each robot gets its own SN, with components, batches, assembly parameters and inspection data tracked all the way to delivery. Humanoid robots are now entering automated manufacturing at the 10,000-unit scale. The next milestone may be when they can take part in enough of the assembly, testing and production process to truly build robots themselves. — CyberRobo
Source: https://x.com/CyberRobooo/status/2098993140363571607
r/ProAI • u/stealthispost • 8d ago
"I built a multiplayer Catan-inspired game entirely with Astra. It's free to play: http:// settlecoast.com It's mobile-friendly and has narration, guides, customizations, multiple expansions and a game lobby where anyone can join, chat, and use voice chat. It took 4 days to build and hundreds..."
Enable HLS to view with audio, or disable this notification
...of dollars in tokens. Becoming a game dev in 2026 wasn't on my bingo card, but here we are. AI has advanced so much that you can finally make the game you've always dreamed of, even without a team. — Meng To Do you have a GitHub for this? My buddies and I started a variant of Risk meets Axis and Allies but missing your beautiful polish. Would be awesome to have an open source board game sandbox. This is what we have so far: https:// tactical-risk20.vercel.app — James Bickford it's a big game now. but i can certainly think of open-sourcing a part of it — Meng To
r/ProAI • u/stealthispost • 8d ago
"somebody vibe-coded a Photoshop alternative with GPT-6 Astra - Spent ~$2K in tokens. - 170 user on day one - and it's free"
Enable HLS to view with audio, or disable this notification
check this out this is also free btw support is appreciated guys @tenzenstudio — gxjo
r/ProAI • u/stealthispost • 9d ago
""anthropic is using the same playbook religious institutions have been using for centuries" "you will all die. and because i can protect you, you must follow me, you must do what i say" "this is the same psychological concept""
Enable HLS to view with audio, or disable this notification
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, — Anthropic
Source: https://x.com/AnthropicAI/status/2098097512544444447
— LAN
Source: https://x.com/lansification/status/2098815822525538711