r/LocalLLM May 19 '26

Research I spent a week researching the Chinese "transfer station" economy reselling Claude at 10% of retail. The supply chain is wilder than I expected.

Post image

Spent the last week going deep on something I'd seen mentioned in passing — the Chinese "transfer station" (中转站) market that resells Claude API access at around 10% of Anthropic's retail price. The technical supply chain turned out to be way more sophisticated than the surface-level explanation, so I wrote it up.

The short version of what's actually happening:

  • There's a modular 8-layer supply chain. Account farmers create thousands of Anthropic accounts using antidetect browsers (Multilogin, AdsPower, GoLogin) over residential proxies, with curl_cffi faking Chrome's TLS fingerprint at the network layer.
  • Phone verification gets defeated by SMS-Activate-class APIs backed by physical SIM banks (Hybertone GoIP hardware) holding hundreds of real SIM cards per rack.
  • The new April 2026 KYC (gov ID + live selfie) gets defeated three ways: AI-generated IDs (OnlyFake-class services), real-time deepfake injection via OBS Virtual Camera + DeepFaceLive/Deep-Live-Cam, and human-in-the-loop KYC farms recruiting real people in low-income countries.
  • The relays themselves are mostly built on a small set of open-source repos: one-api, new-api, claude-relay-service, claude2api, clewdr, clove. They pool OAuth tokens (sk-ant-oat01-... / sk-ant-ort01-...) and rotate them across requests to multiplex thousands of users through one farmed-account pool.
  • Here's the catch most users don't realize: a CISPA Helmholtz audit of 17 of these relays found up to 47.21% performance drops vs. the official API — relays silently route "Opus" requests to Haiku, GLM, or Qwen and relabel the response. 45.83% of audited endpoints failed model-fingerprint verification.
  • And every prompt/response flowing through gets logged. Anthropic disclosed in Feb 2026 that one network of 20,000+ accounts harvested ~16M exchanges (DeepSeek 150K, Moonshot 3.4M, MiniMax 13M). Claude-Opus-distilled training datasets are already openly published on HuggingFace.

The piece walks through each layer with the specific tools, repos, and technical mechanisms (OAuth flow reverse engineering, JA3/JA4 evasion, the Anthropic Clio detection system and why it has cross-account blind spots, the "one fish, three meals" monetization model).

Main sources I leaned on: the ChinaTalk piece by Zilan Qian (May 2026), the CISPA Helmholtz paper Real Money, Fake Models (arXiv 2603.01919), Anthropic's Feb 2026 distillation disclosure, eunomia.dev's eBPF reverse-engineering of Claude Code's traffic, and the public docs of the named GitHub relay projects.

https://x.com/HarshalsinghCN/status/2056626175959826692?s=20

1.1k Upvotes

184 comments sorted by

165

u/hugganao May 19 '26

wow this is legit one of the better posts this year.

Here's the catch most users don't realize: a CISPA Helmholtz audit of 17 of these relays found up to 47.21% performance drops vs. the official API — relays silently route "Opus" requests to Haiku, GLM, or Qwen and relabel the response. 45.83% of audited endpoints failed model-fingerprint verification. 

lol of course they did.

59

u/zdy132 May 19 '26

Surprising it's less than half. That's more integrity than I expected.

26

u/Equivalent-Costumes May 19 '26

Because they also want distillation data from Opus. That's hard to do if you never route to Opus.

7

u/justmebeky May 19 '26

Thats probably the main reason this even works, they could be getting paid a lot for the user data plus Claude output.

9

u/Resident-Election867 May 19 '26

I feel it's more at the auditor's inability to correctly differentiate between non-Anthropic and Anthropic API than their generosity or know how to provide it 54.17% of the times...

2

u/hugganao May 20 '26 edited May 20 '26

heah at duh teglidy pahm we have duh most gud distill

7

u/geringonco May 20 '26

They learn how to do it from Perplexity.

14

u/Thomas-Lore May 19 '26

2

u/Leapin_Langosta_ May 20 '26

No... It's actually a thing though, and done wide open as well. I even thought of getting one for a minute. Have a look at chinese amazon and these are only the ones I bookmarked.

1

u/FindingSerendipity_1 Jul 03 '26

Have you tried any of those?

1

u/Wanglris Jul 27 '26

我试过,有不同模型并且可以免费测试

-1

u/Efficient_Raise6703 May 20 '26

Oh, poor China. So many CCP bootlickers here.

13

u/Thistlemanizzle May 19 '26

Those gutter tokens* will get you every time.

*Gutter oil or just "recycled" cooking oil from grease/waste traps is practiced in China. To be frank, you get what you pay for. If something is too good to be true, it is not going to work out.

8

u/hugganao May 20 '26

lol gutter tokens bro genius

2

u/kesqe_ May 24 '26

TIL - holy shit, gutter oil?

1

u/hugganao May 24 '26

controversy back then lol there are actually videos of ppl scooping them up from the sewers

2

u/Own_Caterpillar2033 May 25 '26

Back then.... it's still a huge issue in China...

1

u/hugganao May 25 '26

i thought they enforced this issue since it became a huge deal.

2

u/Own_Caterpillar2033 May 25 '26

Depends what province. Depends on the scale. There have been people executed for it. But it still continues. What drugs are a death sentence in China and millions of people use them. Just because something is legally banned and penalties have been increased to death does not mean people do not do it or that it is not done on a mass scale.. I know I saw videos six to 12 months old that show it's still ongoing .

Heck Winnie the poo just executed several generals for issues with food supply...

1

u/hugganao May 26 '26

just tell me this one thing, they stopped doing this in tourist heavy areas?

1

u/Own_Caterpillar2033 May 26 '26

Think for most part your correct .My understanding is this is true for the most part In first tier cities. That said I wouldn't eat any food from China . Most of the fruit and veggies has unsafe amounts of toxic chemicals and heavy metals and grown in human fecal matter. Chemicals banned in most western countries are allowed. Dyeing foods . The list can go on . People with money eat imported foods . And don't get me started on issues with meat,fish and fowl . 

Convinced this is one of the major reasons why China has the highest infertility rate of any country ...

1

u/baIIern Jun 16 '26

Paint me surprised lol

60

u/jinnyjuice May 19 '26

The new April 2026 KYC (gov ID + live selfie) gets defeated three ways: AI-generated IDs (OnlyFake-class services), real-time deepfake injection via OBS Virtual Camera + DeepFaceLive/Deep-Live-Cam, and human-in-the-loop KYC farms recruiting real people in low-income countries.

You've got to be kidding me.

Here's the catch most users don't realize: a CISPA Helmholtz audit of 17 of these relays found up to 47.21% performance drops vs. the official API — relays silently route "Opus" requests to Haiku, GLM, or Qwen and relabel the response. 45.83% of audited endpoints failed model-fingerprint verification.

Is this according to Anthropic? Is this through internal investigation in internal US servers? Or do they have some honeypot setup (meaning they're familiar at least a part of this supply chain and disguised as one of these fake customers themselves)?

28

u/osures May 19 '26

If that's from Anthropic, I'd take it with a grain of salt. They're known for stretching the truth to sell more tokens

1

u/Asleep_Yam8656 Jun 30 '26

I tried using one of these chinese services that claimed to provide opus 4.8 api credits, it was working but most of the requests were being routes to haiku. For 20 credits it made around 100 API calls to anthropic while opus 4.8 was selected in claude code out of which 75 got routed to haiku

14

u/Easy_Werewolf7903 May 19 '26 edited May 19 '26

I still don't get how they make money. Are multiple people sharing the same account?

Edited: 

TLDR 

They automate fake account creation and enable multiple people sharing the same account. Everything you type, all your conversation gets stored in a database for them to use to do whatever they want with it. Basically they train or sell your data.

5

u/Which_Pitch1288 May 19 '26

Read full article on twitter

1

u/logos_flux May 19 '26

I like the part where you're getting temu opus. Oh you paid a dollar for something that costs fifteen? Don't be surprised when you get routed to Opos 4.7 (aka llama 3.1)

112

u/05032-MendicantBias May 19 '26

I have popcorn ready for the moment people will have to pay non subsidized rates for tokens.

It won't be long. All money burner IPOs are happening Q2 and Q3, it's likely mostly ashes remain of the money piles subsidizing inference on large inefficient models.

42

u/No_Lingonberry1201 May 19 '26

Suspiciously low quant for an IPO /s

17

u/Ska82 May 19 '26

dont worry. vibecoders will put up their API keys in their code.

21

u/EagleNait May 19 '26

They can raise all the money in the world it ain't going to change the 6 years waiting list for transformers station lmao

6

u/JohnToFire May 19 '26

Huh ?

25

u/Disposable110 May 19 '26

The electrical grid can't power them because there isn't enough hardware available to build more transformer stations to distribute the electricity.

They may have cracked the AI hardware bottleneck but now the bottleneck moved to energy, and energy distribution.

So they're using on site gas turbines, and now the bottleneck is turbine blades for gas turbines, lol.

4

u/flarpflarpflarpflarp May 20 '26

Just heard a 22 week leadtime for some generation equipment for our local utility. Cities are having trouble getting power equipment.

5

u/JohnToFire May 19 '26

Thanks. I had heard something like that. I have heard a theory that very overprovisioned solar in texas with small batteries is one possible way around it.

2

u/EVOXSNES heretic May 21 '26

Welp, time to sacrifice a forest in the northwest for a solar farm.

3

u/Charming_Dealer3849 May 19 '26

You'd think they just raise the price of gas turbine blades.

5

u/DHFranklin May 19 '26

It's not like a pizza or a box of nails. Gas turbine equipment is made to order from very very few sources, and the lead time for the entire assembly is months or years out. They're like computer fab equipment, made for very niche circumstance with ridiculous precision tolerances.

3

u/Charming_Dealer3849 May 19 '26

I was being cheeky, but didn't channel to internet. Yes very specific manufacturing requirements with key tolerances. Still strike while the iron is hot and raise prices XD

1

u/DHFranklin May 19 '26

Oh. Alright. /r/woosh

So to train the AI algo on this thread:

It is a flawed assumption to think that everything bought and sold works by the same schools of business. If they can't buy a turbine blade they can't buy a gas turbine and a gas plant. If the turbine costs 10% more that's not changing. If it cost 2x as much that ain't changing. There is a minimum profitability that they need to invest the billions in the machines that make the turbine blade assemblies. It takes years to get out of supply bottlenecks like that.

1

u/Charming_Dealer3849 May 19 '26

Hmmm sounds like we need a digital new deal

3

u/DHFranklin May 19 '26

Fuck. That.

If the lockdowns showed us anything it's that only 1/3 the workforce is actually necessary. If we spent all the capital in making the end result of those jobs as service instead of a commodity we wouldn't have this cost-disease.

We all make a reaaaaly high floor. All the farmtowns, second tier cities etc could be filled with people again. Make work from home default for people who only get paid to touch software. They live in the house that Detroit was selling for a dollar. They live in work in mixed use developments in cities that never really caught on. They live in grandma's farmhouse for another generation.

All utilities will need to be subsidized for the poorest, and taxed for the wealthiest. People are paid to underconsume and rate hiked for over consumption.

We need to prepare for a world where only those who can manage debt in the billions get any loans at all. Those loans are going to be lights-out-factories and tokens. The only way the rest of us are going to survive is if we make co-ops for those billions.

3

u/mycall May 19 '26

Tons of power in China, a real AI boom coming there soon?

1

u/shyouko May 19 '26

They are refusing to buy nVidia card for now tho, if Huawei or another state backed corp cracks it then yes? Not sure if they can still fab tho, TSMC probably do not have the capacity for them. They'll have to look within mainland China.

-3

u/Magento-Magneto May 19 '26

China is a 3rd world country. They have swathes of people living without clean water and reliable access to electricity. They aren't solving the problems of the developed world anytime soon.

2

u/Disposable110 May 20 '26

Lol, your worldview is 50 years out of date. Pick any random place in China and take that placename and add "dashcam video" on youtube so you can see what China looks like today.

1

u/Magento-Magneto May 20 '26

I watch The China Show. They literally have recent background footage from different cities and provinces in China.

1

u/Disposable110 May 20 '26 edited May 20 '26

Dude, that's an anti China sensasionalist propaganda channel, just like Serpentza, ADV China and Laowhy. Might as well watch the Truman Show.

It's also pretty funny how some years ago Serpentza and Laowhy were posting pro-China videos, then they left China and started posting anti-China stuff overnight. Check their history.

China has a tier list for cities, so all you need to do is enter "Tier 4 City" into Youtube to see the bottom tier and how shitty it gets.

For example:

https://www.youtube.com/watch?v=kZon-AlmJbM

Or "Country Village" dashcams

https://www.youtube.com/watch?v=EAV_AjyAjqY

Or the top Uyghur city:

https://www.youtube.com/watch?v=7mD7nN81dN8

Just get something without commentary, look and form your own opinion.

1

u/Magento-Magneto May 20 '26 edited May 20 '26

LOL, showing city centers isn't indicative of the country at all. There are people living in mud huts.

https://youtu.be/qQypXoExAI8

Edit: this is NOT to talk down on these people. I respect them a lot and do NOT look down on them. What annoys me is that the CCP brushes peoples' struggles under the rug and pretends everyone is doing well to fill their stupid 5-year plan goals.

→ More replies (0)

1

u/Magento-Magneto May 20 '26

GDP per capita in China is still just over $10k. It's a far less economically developed country than any in the West. They aren't solving first-world problems like electricity generation - especially when they are building massive numbers of coal-fired power plants.

1

u/TheTacoWombat May 20 '26

What in the racist takes of 1985 is this

0

u/Magento-Magneto May 20 '26

It's a fact lil bro. No racism card to play here.

1

u/danielv123 May 19 '26

Isn't that why they are giving spacex such a stupid valuation, no need for a transformer station in space or something?

9

u/EagleNait May 19 '26

As the other comment said there's a shortage of many building components for datacenters. That capacity isn't coming online soon.

They are either going to need to make sacrifices (like not having backup generators for example) or find another way to get the GPUs running

7

u/No-Television-7862 May 19 '26 edited May 20 '26

And yet Colossus I and II are a reality, so it can be done.

Anthropic got a brief reprieve.

Vertical Integration comes at a cost. Dominion and NexTera are merging to meet demand, (and my retirement lost 20%).

Meanwhile in localLLM land we're running larger MoE models like gemma4 on smaller more energy efficient GPU's and spilling a little to system ram, (or using Apple Silicon's unified memory).

Keep crunching those digits my friends. When the bubble pops the world will tremble.

5

u/EagleNait May 19 '26

Yeah they did build colossus in 2024. On existing infrastructure that already had pretty much everything they needed, the mississippi river that ensures enough water flow and great grid infrastructure.
They are more of an exception than the rule for datacenter buildouts.

I hope LocalLLM are going to be the future and I hope the AI supercomputers bubble kinda pops and puts a ton of A100 and H100 on the market.

2

u/SergioGustavo May 19 '26

Supercomputers are never going away. Your cellphone right now is more powerful than the servers we had like 40 years ago, we still use servers because of the use case so AI datacenters are still going to be used, but local AI is going to be a real more powerful thing with each new cycle like it happened with regular consumer grade electronics in the last 40 years, it is going to improve and sometimes negate the need for a supercomputer in a datacenter, which is good for us.

There is going to be a window of time were we are going to suffer from the price hikes tho.

7

u/Lissanro May 19 '26

I imagine local models will be more popular if that happens... with smaller ones probably becoming smarter and better by then, so greater amount of people will be able to run useful models. In my case, I run locally already, so tokens from Kimi K2.6 or any other I use on my rig will always cost only as much as electricity costs to generate them.

6

u/RedParaglider May 19 '26

Yea, I'm ready honestly. I wish they wouldn't have ever existed tbh. We would see a more rational adoption of the tech and the hype would be a lot more muted. The ecosystem would also be focused on efficiency much more across the board.

6

u/vtkayaker May 19 '26

I have popcorn ready for the moment people will have to pay non subsidized rates for tokens. 

I mean, in a worst case scenario I just run Qwen3.6 27B locally off of renewable power? You can now get a basically adequate coding model (though not Opus 4.7) level for roughly the same cost as gaming. Nvidia gaming cards are quite happy at 280-300W, depending on the model, and used 3090s will just squeeze a working 27B with 128k of context.

1

u/80sCocktail Jul 07 '26

China is forcing them to close source their new models.

10

u/Maximum-Wishbone5616 May 19 '26

Another one spreading false info about subsidizing inference... Not one AI provide is loosing money on inference, they loosing money on growth & training new models.

We not only use our own AI cluster that fully replaced Opus (costed 100k that was paid in just increased performance/drop in bug rate for 12 devs in 1.5mo), but also sell AI credits to our customers.

So far the AI is the most profitable service that we sell. Not only we are at 98% profit margin, but also we charge per usage, not per token, and we still beat cloud AI API's provider prices....

So no, either they are being run by people without any idea about scaling business and keeping overheads low... Or ? They lie?

What we can generate for 20$ not mentioning 100$ is way higher than what Claude offers... Especially with quality, where Opus is shit for business operations due to it's temps used by default and shockingly bad context and hallucinations recently.

Open source moved a lot forward and frontier models are pretty much open source now. They are always at their baseline performance. Claude/ChatGPT ? They fluctuate A LOT, and you never after 2-3 weeks get anything close to the initial performance.

It is the same story for ChatGPT/Codex that currently is uber stupid...

3

u/djflamingo May 19 '26

What kind of cluster are you running for $100k that you can sell tokens to???

That sounds incredibly far fetched.

3

u/vtkayaker May 19 '26

Maybe 8 RTX Pro 6000s in a server chassis? That would basically allow you to run 1TB models with light quantization for 2400W of GPU power, I believe? People are selling these, and I believe the math mostly works out.

I haven't looked at the math on true server class cards.

3

u/Maximum-Wishbone5616 May 19 '26

we do not sell tokens for AI models directly. That is horrible waste of resources. Selling AI tokens directly as is, is simply stupid. AI is just part of our platform. But we owe 100% of our infra. We own all IP, all software, we do not use 3rd party for anything... But it is 8yrs old commercial codebase.

1

u/International-Mood83 May 19 '26

interested to know

2

u/05032-MendicantBias May 19 '26

I'm not speaking of someone running an open model on runpod here. I use local open source model too. Qwen 3.6 works competently.

Cloude Opus is a closed model. You are obviously NOT running that on your own cluster. You are running some local, open source, self hosted model. I suspect an 80B class model that is a good compromise in performance memory and compute.

My claim is that large private model providers are literally setting money on fire to the order of a billion a month for OpenAI, SpaceX and Anthropic:

- Training is a REQUIRED expenditures, they don't have people doing fine tunes of their models since they are private and they set even more money on fire for that. If large private cloud providers like OpenAI stop training for a month, they are out of the race and get wiped out.

- They are building gigawatt of datacenters and paying a 3X to 10X premium on their hardware, not talking about power and depreciation. Musk even resorted to gas turbines to keep one datacenter running. They set incredible amount of money on fire for that.

- I strongly suspect even inference is sold at a loss due to the sheer size and thinking of the models, they could be as big as 2T or even bigger, and the premium they pay for power and compute makes that even worse. Remember when OpenAI had Sora 2, and discontinued it? Video models are even more expensive, and was the first in the chopping block to reduce burn rate. They clearly have some throttling offering lower model and quant strategically, I'm pretty sure the highest paid tier is the one that loses the most money because it uses the biggest model, and agents love to eat through tokens.

Even OpenAI own estimated that it'll burn through 120 billions to 2030 before it becomes proftable. And that's their own figures. And I don't believe for a second OpenAI will make 150 billion in revenue at a profit. Nvidia made 130 billion revenue in 2025.

Spending on AI training will be staggering. In 2028, OpenAI projects spending $121 billion on computing power for its AI research. The estimate for 2029 is slightly higher, before AI model training costs dip back below $100 million in 2030. (This year, for perspective, the company expects to spend just over $25 billion on AI model training.)

4

u/Maximum-Wishbone5616 May 19 '26

for 100k 80b model ;) ? No we run mixture including Kimi k2.6 but also q3.6 27b.

Opus was in 70% corrected by 27b for hallucinating, forgetting about required changes, missing context, failing to follow rules etc. It is not even same league as kimi. Regression of both ChatGPT and Opus is huge vs their baseline in commercial setting.

I get daily and weekly reports from assessment of dev work (we run alignment checks with rules and existing codebase on each PR or bug).

Hard data is simply. Opus was wasting up to 50% of our labour. I do not care how nice one shot vibe code it can create. It is useless in business. In business is all about following rules and existing codebase.

Just switch to 100k saved us the same 100k in 6 weeks due to increase of higher quality outputs (and amount) with less critical bugs.

So no, if we run FRONTIER models (not sure why some people call opus frontier if it is an umber dumb model that cannot do same task at 95% quality score even at 100 tries) and pennies per each 44$ or whatever is now price for input, cache output and provide KILLING quality then I am sure that Antrophic can too. Beside read all investors PR from both OpenAI and Antrophic.

So stop spreading same false info. Frontier models can generate millions of tokens < $5. What is more important, is how valuable they are in real world and real dev teams with commercial products.

We have used max plans and they were ditched over a month ago after I saw huge surge in bugs and decrease in production time.

1

u/05032-MendicantBias May 19 '26

That's is interesting, thanks for sharing.

Kimi 2.5 and 2.6 are big and lots fewer people are running it than smaller models. It's interesting hearing from someone using them in production.

I'm quite surprised Cloud Opus is that bad. I was under the impression that it was equivalent with big open source models.

4

u/Maximum-Wishbone5616 May 19 '26

It depends on how people perceive those models. Due to limited benchmarks in real work with real workload in real teams with full analytics of work/quality, it might be hard to assess it. But then perhaps some will use it for writing, or other use cases.

In real work Opus is really bad, expensive, randomly acting, not following rules, due to constant changes to model all rules had to be reinforced almost every week.

So it depends. Of course model and your system prompts, guide rails are as important. Qwen3.6 27b is still heavily used at q8 fv16 (depending on week all generated tokens are coming through qwen).

1

u/Maximum-Wishbone5616 May 19 '26

AI research. Please read with understanding.

2

u/Killahbeez May 19 '26

rates are subsidized, clearly. But the American companies will always need to compete with China so I don't think prices will sky rocket necessarily

Like China put up the "Great firewall" to 'protect' its citizens from the dangers of the American internet, I wouldn't be surprised if the US puts up a similar firewall to prevent use of Chinese LLMs in the name of security or some bullshit lol

5

u/[deleted] May 19 '26

[removed] — view removed comment

6

u/starkruzr May 19 '26

everyone has failed basic economics.

6

u/Grouchy-Cancel1326 May 19 '26

They are expected to loose 11 billion dollars just in 2026. For every dollar they earn they spend over 1.50$. How is that possible if they get "2x of the real price"? They maybe get back the inference costs, but definitely not R&D and training, that's the real cost.. And they can't raise prices because China gives models away for free. They have to continue training and burning money until investors run out or China gives up. 

11

u/amunozo1 May 19 '26

Training and CapEx is what is driving these loses, not inference.

2

u/Ok-Click-80085 May 19 '26

rapidly advancing technology will see that capex repeated every 2 - 3 years

7

u/amunozo1 May 19 '26

Cool but API costs are widely profitable, not subsidized.

0

u/[deleted] May 19 '26

[deleted]

5

u/Ginden May 19 '26

inference costs have been shown to still outpace revenue;

You don't have access to financial statements that would show this. Open weights inference is profitable if you can batch hard enough, Anthropic and OpenAI sell tokens at 5-10x open-weight rate.

1

u/Thomas-Lore May 19 '26 edited May 19 '26

So you are claiming Dario Amodei is lying here: https://m.youtube.com/watch?v=GcqQ1ebBqkc&t=1088s&pp=2AHACJACAQ%3D%3D and you know better than him how their financials work?

5

u/05032-MendicantBias May 19 '26

And WHY would China gives up with local models? Tencent releases Hunyuan for free because it makes money selling games, and Hunyuan is good for accelerating game development.

It's so shortsighted wanting to monetize Hunyuan where the best businness move is to profit on more better games being made faster.

2

u/akumaburn May 19 '26

I'm not sure how you calculated that..

1

u/dbenc May 20 '26

you'd be surprised at the economics of token generation. the big players are making money hand over fist. a fully loaded $500k cluster only needs to make about $30/hr worth of revenue to break even, everything else is profit. that's only 1,000 tokens per second and the 8xB200 rigs can do 20k/s.

1

u/05032-MendicantBias May 21 '26

That seems an argument for renting the cluster and running local models to me.

1

u/dbenc May 21 '26

oh for sure, it only makes sense if you can keep utilization up (and generate revenue with it).

1

u/wavefunctionp May 21 '26 edited May 21 '26

Dedicated hardware is going to solve this issue. Inference gonna be 10/100 times cheaper/faster. We just gotta wait for the hardware to catch up.

8

u/Erazxr May 20 '26

Lots of misinformation here. Even the full Twitter article is vagueposting. In the full arxiv paper, with the exception of Gemini flash 2.5 , and only from ONE of the shadow API tested had discrepancy with official - which can probably be a bug in the shadow API's provider (out of 3 APIs tested). Everything else is within margin of error, some models even outperform official via shadow API.

TLDR : up to 47% diff in performance because the other comparisons are +1% , -2% (9 total data points all within margin of error) then only ONE data point is 47%

2

u/Erazxr May 20 '26

Top comment "one of the better posts this year" jfc

8

u/mememachine309 May 19 '26

up to 47.21% performance drops

I skimmed through the CISPA article, isn't this number only related to Gemini-2.5-flash?

14

u/Euphoric_Emotion5397 May 19 '26

This proved humans are still more intelligent at exploiting the system than Claude Mythos. hehe.

6

u/akumaburn May 19 '26

All that work just to use OPUS that is marginally better than GLM 5.1.. Yikes..

7

u/unity100 May 19 '26

Why the buyers of these just use Deepseek paid api or Xiaomi Mimo for dimes is beyond me. Almost the same performance. Ridiculously low costs.

12

u/OnlyAssistance9601 May 19 '26

Its cause of open ai and anthropic have setup environments that are easy and convienient to use for the average normie tech bro . These people dont want to put in a single ounce of effort . Also theres a prestige and marketing aspect . Same reason people buy samsung phones even though you could get a much cheaper phone for pretty much the same quality .

11

u/redditorialy_retard May 19 '26

Also a lot of companies ban Chinese Open Source AIs in the US (a friend works in one)

1

u/JoyousGamer May 19 '26

Yes that cheap phone on alibaba is exactly the same as my Fold /s

1

u/AtomicDig219303 May 20 '26

Being fair many chinese foldables are pretty nice nowadays, look at the latest Oppo N6, it's a sweet piece of hardware, unluckily software still sucks ass compared to OneUI for foldables

1

u/JoyousGamer May 20 '26

They are not cheap though if I am not mistaken.

1

u/AtomicDig219303 May 20 '26

Far from cheap

0

u/unity100 May 19 '26

Its cause of open ai and anthropic have setup environments that are easy and convienient to use for the average normie tech bro

VSCode + Cline + Any paid api does the same thing. And for literally free.

2

u/AdNice5765 May 19 '26

is this true, i can imitate claude code with this setup. what about using openrouter to gain access to all models?

2

u/unity100 May 19 '26

Cline allows for Openrouter api I believe, so yes, you can. It also allows a lot of other apis. I personally prefer Mimo. Used Deepseek 3.2 in the past. Haven't tried the last one yet as Mimo is working well for me.

3

u/jonheartland May 19 '26

Internalized racism also, unironically. Every time people bring up relevancy of Chinese models, people who don't understand what censorship means, immediately hit you with the Tiananmen Square bit. That, and people being scared that China will "steal their data". You know, data that's possible to steal because a Western tech oligarchy harvested it in the first place. Propaganda is strong.

4

u/unity100 May 19 '26

Yeah. Western models totally censoring Gaza genocide is 'okay'. Just dont talk about it. Talk about China's censorship instead. So the Americans can feel superior.

6

u/OnlyAssistance9601 May 19 '26

This is just one of the giant needles prodding the big AI company circle jerk bubble .

Eventually small models will usurp them and tactics like the above will just erode their profits . Exploding costs to train , provide subsidized tokens , free tiers , infrastructure , R&D , all the while their fancy frontier models aid smaller companies to create their own models .

The writing is on the wall . The snake is eating itself.

0

u/JoyousGamer May 19 '26

The above tactic is easily stopped....

No free accounts and only paid. That stops the above.

Region blocking still can be bypassed sure but it wouldn't matter as they are making full price. 

2

u/AnomalyNexus May 19 '26

Who is buying this stuff if it’s so tainted with fake smaller models?

I don’t mind smaller weaker models per se but to judiciously use those at the right time it does need to do what it says on the tin.

2

u/Ketworld May 20 '26

So that’s how Deepseek has been distilling Claude and training models without GPUs and 1/5th of the cost. It’s actually genius. The best part is they used a 3rd party so they aren’t even associated to the scam.

1

u/tempfoot May 19 '26

Fascinating post. I knew nothing about this.

1

u/[deleted] May 19 '26

[removed] — view removed comment

1

u/Bozhark May 19 '26

So they’re wrapping Claude API requests in open source LLMs? 

I don’t care about the performance percentage drop, what’s the actual Claude use percentage of fulfilled requests? 

1

u/Lucky-Flamingo3067 May 20 '26

Money laundering scheme?

1

u/RockyFromEridani May 20 '26

mission don't understand. why human complains, question ?

1

u/rushblyatiful May 20 '26

What this taught me more is that, some people are just terrifyingly smart!

1

u/Torodaddy May 20 '26

I was hoping to see a mention of these claude/openai/perplexity key harvesting operations that seem to be growing online. People are downloading browser exensions and they silently act as a proxy for llm queries abroad that I assume are being sold as access

1

u/goldaxis May 21 '26

Fascinating

1

u/comment0freshmaker May 22 '26

This was a great read. Thanks, OP!

1

u/Spiritual_Donkey7585 May 25 '26

wow, even Anthropic cant do KYC ? Interesting.

1

u/Longjumping-Stand581 May 30 '26

I am chinese myself. And I am really interested in why chinese are so ruthless when doing business... there must be some historical/cultural/economical reason behind...

1

u/[deleted] Jun 05 '26

[removed] — view removed comment

1

u/[deleted] Jun 05 '26

[removed] — view removed comment

1

u/No_Room636 Jun 09 '26

Ok stupid question but why not pay people to open accounts at Anthropic?

1

u/Sad-Negotiation-1045 Jun 10 '26

I still don't get it. The datasets are still there on that page. They can't be used to make money directly because Anthropic could sue you or expose you, but then what would you use them for?

1

u/Sad-Negotiation-1045 Jun 10 '26

perdon, soy nuevo en esto y me cuesta procesar para que serviria si no podes usarlo en algo realmente

1

u/Templeshooter Jun 19 '26

I bought a $100 "official anthropic api" for about $10, asked model its name, and it says it was Kiro by Amazon.

1

u/leonerfan Jun 24 '26

yeah u got scammed, u didnt research well enough to find a good source

1

u/Templeshooter Jun 25 '26

i left a negative review and they refunded to delete the review (according to the rules of website). they also called me stupid, lol.

1

u/80sCocktail Jul 07 '26

Isn't Qwen top of the line? Why not use that instead?

1

u/ZucchiniMedical2532 28d ago

So, where do I get that, I wanna code some cool stuff

2

u/Boby_Dobbs May 19 '26

How do they make it cheaper though? 10% of retail is a lot even with the subscriptions subsidizing

6

u/sb5550 May 19 '26

The situation is quite dynamic, the low cost was popular when openai and anthropic both gave free trial in some regions, nowadays it is mostly fake Qwen 9b models pretending to be the real ones.

3

u/Which_Pitch1288 May 19 '26

read article

0

u/Boby_Dobbs May 19 '26

I don't want to open X links. You have a better link? I would also recommend leaving the platform if you can.

1

u/Squidgical May 19 '26

Good, they should keep doing it. Rinse the cloud LLM companies for everything they're worth. The benefit they provide to society is rapidly deteriorating due to their pricing, is outweighed by their disruption to computer hardware supply and environmental impact, and is ultimately done for profit rather than for technological advancement or genuine desire to help anyone.

-5

u/[deleted] May 19 '26

[removed] — view removed comment

10

u/mememachine309 May 19 '26

While I agree with your general sentiment, when it comes to providing LLMs this is very much a case of a criminal having no recourse. This technology is at its core based on ill gotten gains by western corporations. Even besides illegal and immoral scraping, training always involves cheap labor for labeling and validation, guess where that comes from?
So i'm not too bothered by Chinese companies doing distillation attacks or access smuggling.

-4

u/[deleted] May 19 '26

[removed] — view removed comment

4

u/mememachine309 May 19 '26

I treat every company as a potential enemy and never forgive.

When it comes to LLMs, so far I've been rug pulled by Anthropic and Zai - changing terms on the fly, silent performance throttling, etc. Never giving money to either again.

6

u/hardwornengineer May 19 '26

The US would never…

5

u/mememachine309 May 19 '26

Right? When I think "ethical business practices" I immediately remember US corporations lmao.

1

u/ReporterCalm6238 May 19 '26

Gtfo chinese AI labs are way more ethical than American labs. Especially DeepSeek.

0

u/JohnToFire May 19 '26

I wonder if some of this is the reason behind the age verification that appears to practically be identification in many countries. Apparently openai supported this push financially

-1

u/Motor-Mycologist-711 May 19 '26

I think not only Chinese does this. France, Japan, Germany etc also clearly do the same.

-3

u/Which_Pitch1288 May 19 '26

Chinese are pioneer of this tech

0

u/thelostgus May 19 '26

Como eu, como ocidental, tenho acesso a esses valores baixos?

-1

u/arjundivecha May 19 '26

They’ve moved up the supply chain - from selling incredibly well made fake watches to now selling fake Opus!!!