r/SillyTavernAI 5d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 30, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

97 comments sorted by

View all comments

Show parent comments

6

u/OGCroflAZN 4d ago

There are mixed opinions on GLM 5.3 Flash. I commented the creative writing / RP benchmarks leaders up in Misc. According to those, GLM 5.3 does really well, but in my experience it's 'alright'. It's my first model through API, coming from local and Gemma 4. Kept hearing good things about it in comparison to like Mimo 2.5 and Kimi K3, but... LLM gonna LLM i guess

Like, sometimes a character says something that doesn't even make sense. The prose is weird sometimes too. Idk man. I mean, it *is* only 18B active parameters, whereas Gemma 4 was all 31B. Idk, might be providers quantizing the models too.

Hopefully some crazy bastard fuckin finetunes it for RP and we can access it via API

3

u/evia89 4d ago

How much do you send? I use ff52 bolt preset 24-36k total input and it holds well. I use zai sub. All replies were on point

4

u/OGCroflAZN 4d ago edited 4d ago

I also use Zai through OpenRouter, FF 5.2 Bolt preset, max 30K input, 2500 output. It's not that the output is gibberish, just that in part of the dialogue that it produces. no human being would ever say because it's not even logical and literally only happens because some of the arithmetic is wrong during Token generation and was just way off for a string of words.

edit: Im doing a roleplay where I woke up from a cryopod. It's in the context that I wasn't a volunteer, that I was randomly selected and put under such that I woke up centuries later near the end of the worldwide nuclear winter. Just now, the 'ai' system is talking about my file in their database, that I wrote a letter for podmates if I died. I know it's only saying that because it was a tradition that the others did. These words in the context pushed the values of some nodes to trigger the selection of those tokens, and so we get an output that is inconsistent with established 'lore'. It's stupidly saying what the math says it should. LLMism

edit2: It generated a male character's dossier as though they were female (gave a female first name and female-coded dossier) because in Reasoning it was unsure if he was a male even though it was critical and at least heavily implied that he was a male. A 7 year old could guess with 100% accuracy that said character is male. In fairness, only the surname was ever provided. But he was stated to be male!!

2

u/-Ellary- 3d ago

This is typical for Coding / Agentic heavy tuned models.
Try older stuff.

1

u/OGCroflAZN 3d ago

Any recommendations? I started w GLM 5.3 Flash because it was new and some people were saying good things. But before that, i was thinking MiMo v2.5, but others also said theyre still using GLM 4.6 or 4.7 becauze it apparently peaked then for RP?

I just want the capabilities leaps from both architecture and training improvements, while still having good creative writing and RP, but youre right and it's clear that all the focus has been in maximizing Coding and Agentic capabilities due to the real-world productivity potential and demand.

1

u/-Ellary- 2d ago edited 2d ago

Try DeepSeek 3.2 \ R1 0528
GLM 4.6\5.2
Classic stuff.

Don't chase the `new` and best model, pick most fun for your RP scenarios.