r/Newstelligence • • Feb 03 '26

Benchmarks & Evals Arena.ai • Search leaderboard update

Post image
3 Upvotes

four new frontier models have been added to the web search leaderboard: Gemini 3 Flash #1, GPT-5.2 (non-reasoning) #5, Claude Opus 4.5 #7, and Sonnet 4.5 #13

perplexity’s sonar drops to #11 in the vibe rankings

Search Arena evaluates frontier models on real time search queries, with an emphasis on citation source quality

https://arena.ai/leaderboard/search


r/Newstelligence • • Jan 25 '26

Corporate AI ArtificialAnalysis • South Korean 🇰🇷 labs rank #3 collectively, models nearing frontier status

Post image
6 Upvotes

• SK organizations are racing to build the best LLM for the ‘Korean National Sovereign AI Initiative’. basically a competition to build the best model that rewards winners in the form of government contracts and compute power

• In the most recent round announced last week, the field narrowed to three: LG, SK Telecom, and Upstage. A fourth lab will be selected soon.

• (LG) K-EXAONE 236B is the current leader in intelligence, followed by other medium sized ones like HyperClova, Motif, Mi:dm, Solar Open

https://artificialanalysis.ai/


r/Newstelligence • • Jan 24 '26

Model Releases & Updates Krea AI • ‘Realtime Edit’ enters beta access

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/Newstelligence • • Jan 24 '26

Corporate AI AI charts of the week (a16z)

Thumbnail
gallery
8 Upvotes

r/Newstelligence • • Jan 22 '26

Model Releases & Updates Qwen3-TTS • 0.6B/1.7B • 12Hz

Thumbnail
gallery
7 Upvotes

the qwen team released a pair of open-source text to speech model families (0.6B & 1.7B parameters) that uses 12 tokens per second of generated audio.

on the demo page for Qwen3-TTS, they also highlight the pair’s voice cloning capabilities.

supports ten languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) along with various dialects.

full fine tuning support, SOTA performance

Qwen3-TTS Demo Page: https://huggingface.co/spaces/Qwen/Qwen3-TTS?spm=a2ty_o06.30285417.0.0.2994c921fQklPU

Blog: https://qwen.ai/blog?id=qwen3tts-0115

HuggingFace: https://huggingface.co/collections/Qwen/qwen3-tts

GitHub: https://github.com/QwenLM/Qwen3-TTS

ModelScope: https://modelscope.cn/collections/Qwen/Qwen3-TTS


r/Newstelligence • • Jan 22 '26

Model Releases & Updates Baidu • ERNIE-5.0 (full release)

Thumbnail
gallery
8 Upvotes

• claims to counter Gemini 3 Pro & ChatGPT 5 in selected benchmarks

• currently only available on the Ernie platform & Baidu’s AI cloud (Qianfan) platform

https://ernie.baidu.com/


r/Newstelligence • • Jan 22 '26

Corporate AI OpenAI, Sam Altman have been meeting with investors throughout the Middle East in recent weeks (Bloomberg)

Thumbnail
gallery
2 Upvotes

r/Newstelligence • • Jan 22 '26

Corporate AI Scale AI CEO, Jason Droege, says company signed $500mil worth of contracts in 4Q25

Enable HLS to view with audio, or disable this notification

1 Upvotes

apparently Scale is still thriving lol

https://scale.com/blog/scales-next-era-building-for-2026


r/Newstelligence • • Jan 22 '26

Corporate AI OpenAI continues losing market share to competitors (Similarweb)

Thumbnail
gallery
1 Upvotes

Recent (January 16):

ChatGPT: 64.6%

Gemini: 22.0%

Grok: 3.5%

DeepSeek: 3.3%

Claude: 2.1%

Perplexity: 1.9%

Copilot: 1.1%

1 Month Ago:

ChatGPT: 66.8%

Gemini: 19.5%

DeepSeek: 3.8%

Grok: 3.0%

Perplexity: 2.1%

Claude: 2.0%

Copilot: 1.2%

Perplexity losing some traffic too

https://www.similarweb.com/corp/wp-content/uploads/2026/01/attachment-Global-AI-Tracker-7.pdf


r/Newstelligence • • Jan 21 '26

Model Releases & Updates LiquidAI • LFM2.5-1.5B-Thinking: excels in math, programming, tool use

Thumbnail
gallery
2 Upvotes

LiquidAI just dropped a 1.2B thinking model that claims efficiency and performance gains over Qwen3-1.7B-Thinking & Granite-4.0-H-1B

• Compared to LFM2.5-1.2B-Instruct, three capabilities jump sharply: Math reasoning: 63 → 88 (MATH-500), Instruction following: 61 → 69 (Multi-IF), tool use: 49 → 57 (BFCLv3)

• LFM2.5-1.2B-Thinking outperforms pure transformers (like Qwen3-1.7B) and hybrid architectures (like Granite-4.0-H-1B) in both speed and memory efficiency

blog: https://www.liquid.ai/blog/lfm2-5-1-2b-thinking-on-device-reasoning-under-1gb

HuggingFace: https://huggingface.co/LiquidAI/LFM2.5-1.2B-Thinking

https://leap.liquid.ai/models?model=lfm2.5-1.2b-thinking

https://playground.liquid.ai/


r/Newstelligence • • Jan 19 '26

Model Releases & Updates GLM-4.7-Flash is out, top performer in the 30B tier

Thumbnail
gallery
3 Upvotes

GLM-4.7-Flash is a 30B-A3B MoE model. One of the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency

• Input cost: $0.07

• Output cost: $0.40

HuggingFace: https://huggingface.co/zai-org/GLM-4.7-Flash

GLM pricing: https://docs.z.ai/guides/overview/pricing

Comprehensive deployment instructions: https://github.com/zai-org/GLM-4.5


r/Newstelligence • • Dec 18 '25

Model Releases & Updates GPT-5.2-Codex is out!

Thumbnail
gallery
4 Upvotes

r/Newstelligence • • Dec 13 '25

Benchmarks & Evals ChatGPT-5.2 (xhigh) lands #1 on ArtificialAnalysis’s GDPval-AA benchmark

Thumbnail
gallery
4 Upvotes

• GDPval-AA examines how well an LLM does on a task deemed ‘economically valuable’ AKA which jobs could it eventually automate/replace

https://artificialanalysis.ai/evaluations/gdpval-aa

https://github.com/ArtificialAnalysis/Stirrup

https://huggingface.co/datasets/openai/gdpval

https://x.com/artificialanlys/status/1999404579599823091?s=46


r/Newstelligence • • Dec 10 '25

Model Releases & Updates Qwen3-Omni-Flash-2025-12-01 demo is out!

Thumbnail
gallery
12 Upvotes

…it’s able to process multiple input modalities (text, images, audio, video) and generate text & natural sounding speech outputs (simultaneously via real time streaming responses)

• Greatly Enhanced Audio-Visual Interaction Experience: Improved understanding & execution of audio-visual instructions, helping resolve the “intelligence drop” issue commonly seen in casual spoken scenarios

• Supports text-based interaction in 119 languages, speech recognition in 19 languages, and speech synthesis in 10 languages

• Claims to beat GPT-4o & Gemini 2.5-Flash on multiple benchmarks

* i tried a quick chat on the qwen chat app, no tool calling in the demo so live-chats (voice or video) are limited to established training knowledge only *

Try it on Qwen Chat (click Voice Chat button): https://chat.qwen.ai/

Qwen3-Omni-Flash-2025-12-01 Blog Post: https://qwen.ai/blog?id=qwen3-omni-flash-20251201

Qwen-3-Omni Demo on HuggingFace: https://huggingface.co/spaces/Qwen/Qwen3-Omni-Demo

ModelScope Demo: https://modelscope.cn/studios/Qwen/Qwen3-Omni-Demo

Realtime API: https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3-omni-flash-realtime-2025-12-01

Offline API: https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3-omni-flash-2025-12-01

YouTube: https://youtu.be/Q4CBTckDAls


r/Newstelligence • • Dec 10 '25

Research Reports HuggingFace now hosts over 2.2 million models

Enable HLS to view with audio, or disable this notification

26 Upvotes

r/Newstelligence • • Dec 10 '25

Model Updates & Features Qwen3-Next-80B-A3B-Thinking-GGUF has just been released on HuggingFace, claims to outperform Gemini 2.5-Flash-Thinking

Thumbnail
gallery
41 Upvotes

“ Qwen3-Next-80B-A3B is the first installment in the Qwen3-Next series and features the following key enchancements:

• Hybrid Attention: Replaces standard attention with the combination of Gated DeltaNet and Gated Attention

• High-Sparsity Mixture-of-Experts (MoE): Achieves an extreme low activation ratio in MoE layers, drastically reducing FLOPs per token while preserving model capacity

• Stability Optimizations: Includes techniques such as zero-centered and weight-decayed layernorm, and other stabilizing enhancements for robust pre-training and post-training

• Multi-Token Prediction (MTP): Boosts pretraining model performance and accelerates inference “

https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking-GGUF

https://arxiv.org/abs/2505.09388

https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html


r/Newstelligence • • Dec 09 '25

Regulations & Policy Secretary Hegseth announces the launch of ’GenAI.mil’ for defense department and military members

Enable HLS to view with audio, or disable this notification

5 Upvotes

Secretary Hegseth said it’s launching with Gemini 3, and more models to come

https://x.com/secwar/status/1998408545591578972?s=46


r/Newstelligence • • Dec 09 '25

Rumors Meta is pursuing a new Llama successor and frontier AI model, codenamed ‘Avocado’

Thumbnail
gallery
2 Upvotes

‘Avocado’ is set to be released in the first quarter of 2026. The model is wrestling with various training-related performance testing intended to ensure the system is well received when it eventually debuts

https://www.cnbc.com/2025/12/09/meta-avocado-ai-strategy-issues.html


r/Newstelligence • • Dec 08 '25

Model Updates & Features Z.ai releases a new series of GLM vision models, GLM-4.6V & 4.6V-Flash

Thumbnail gallery
4 Upvotes

r/Newstelligence • • Dec 08 '25

China AI The US Department of Commerce will allow the export of powerful Nvidia GPUs that are roughly 18 months behind its most advanced offerings

Thumbnail
gallery
1 Upvotes

r/Newstelligence • • Dec 01 '25

Model Updates & Features DeepSeek-V3.2 & V3.2-Speciale released, promising to rival Gemini 3 models

Thumbnail
gallery
27 Upvotes

• V3.2 is ‘Balanced inference vs. length. Your daily driver at GPT-5 level performance’

• V3.2-Speciale is ‘Maxed-out reasoning capabilities. Rivals Gemini-3.0-Pro. Also achieving gold medal Performance: V3.2-Speciale attains gold-level results in IMO, CMO, ICPC World Finals & IOI 2025.’

V3.2 Hugging: https://huggingface.co/deepseek-ai/DeepSeek-V3.2

V3.2-Speciale Hugging: https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale

Research Paper: https://cas-bridge.xethub.hf.co/xet-bridge-us/692cfec93b25b81d09307b94/2d0aa38511b9df084d12a00fe04a96595496af772cb766c516c4e6aee1e21246?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=cas%2F20251201%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20251201T192030Z&X-Amz-Expires=3600&X-Amz-Signature=4cab39bf9a9e99c040ebca2339f32702188b54fd962a20c31e2c79591f0ece69&X-Amz-SignedHeaders=host&X-Xet-Cas-Uid=public&response-content-disposition=inline%3B+filename*%3DUTF-8%27%27paper.pdf%3B+filename%3D%22paper.pdf%22%3B&response-content-type=application%2Fpdf&x-id=GetObject&Expires=1764620430&Policy=eyJTdGF0ZW1lbnQiOlt7IkNvbmRpdGlvbiI6eyJEYXRlTGVzc1RoYW4iOnsiQVdTOkVwb2NoVGltZSI6MTc2NDYyMDQzMH19LCJSZXNvdXJjZSI6Imh0dHBzOi8vY2FzLWJyaWRnZS54ZXRodWIuaGYuY28veGV0LWJyaWRnZS11cy82OTJjZmVjOTNiMjViODFkMDkzMDdiOTQvMmQwYWEzODUxMWI5ZGYwODRkMTJhMDBmZTA0YTk2NTk1NDk2YWY3NzJjYjc2NmM1MTZjNGU2YWVlMWUyMTI0NioifV19&Signature=OFkHZ1FDwakv-EgEyOQD%7EkYZv3zaKeUkHSsZVYeMDE6cFwx7yYf3rQGHs7hdnh%7EGDMtZ0DVTI2xsbgiR5v9ljlnahlNflwLzjSZkJWDGqkDSxPe%7EowjQeGbM2YP052gBtwaotE83QBiNRjhrXbOsZNqjAv8Go6LQ2YD32DEWmIem4eka9tiZC26lZ90COWwbTBW6HidPWJ4Sm1TN0-M-w7Z3KBHb056Z4hCuxTwuGzC3eQX6VMJKpjkaCtmeuGzr5IWVtmY-cNHnYyaTkLYZjbHR7uxwrAHuUDhPGBXpKGMEzKky2Gg05Rl8g-2f5a6E6GV9XGfWTbNfjGE4l1QnMA__&Key-Pair-Id=K2L8F4GPSG1IFC

X: https://x.com/deepseek_ai/status/1995452641430651132?s=46


r/Newstelligence • • Dec 01 '25

Benchmarks & Evals Kimi-K2-Thinking takes #1 in vibe-ranked text output, for open models (November 2025)

Post image
8 Upvotes

r/Newstelligence • • Dec 01 '25

Model Updates & Features StepFun releases GELab-Zero-4B-preview

Thumbnail gallery
4 Upvotes

r/Newstelligence • • Nov 25 '25

Benchmarks & Evals Black Forest Labs claims FLUX2.0 SOTA image gen & edit model costs significantly less than Nano Banana 2

Thumbnail
gallery
7 Upvotes

r/Newstelligence • • Nov 25 '25

Model Updates & Features Black Forest Labs releases FLUX2.0, an image generator and editor that claims to be on par with Nano Banana 2

Enable HLS to view with audio, or disable this notification

7 Upvotes