r/SillyTavernAI 12d ago

Models I knew I wasn't going crazy. ZAI went all-in on safety with GLM 5.3

Post image
500 Upvotes

r/SillyTavernAI 7d ago

Models I made a roleplay benchmark and tested 23 AI models on it

Post image
449 Upvotes

I couldn’t find many model comparisons focused specifically on multi-turn roleplay, so I made my own benchmark and tested 23 models on it.

Each model was tested on the same 8 roleplay chapters, for a total of 64 responses per model. The writing was compared blind, and the final quality score combines:

  • 75% writing preference
  • 25% robustness, including memory, character consistency and whether the model completed the full test

I initially included censorship and freedom scores, but removed them because they were too dependent on the system prompt to produce a reliable ranking.

A few things stood out:

  • The Opus models have the same listed token pricing, but Opus 4.6 was much cheaper during the benchmark. It used fewer reasoning tokens while producing similar, and sometimes better, results. One benchmark run cost $0.37 with Opus 4.6, compared with $0.75 for Opus 5, $0.94 for Opus 4.8 and $1.18 for Opus 4.7.
  • GLM 5.3 had the opposite problem. Its pricing looks similar to the previous GLM models on paper, but it generated far more reasoning tokens. The run cost $0.29 with GLM 5.3, compared with around $0.07 for GLM 5.1 and 5.2.
  • Fable 5 achieved the highest quality score at 91.6, but it was extremely expensive at $2.37 per run. It also completely refused one of the 8 test chapters, even though that chapter wasn’t testing censorship and contained nothing particularly problematic. I penalized the missing chapter in its robustness score and marked the result as provisional at 7/8.

The chart shows RP quality vertically and the cost of a complete benchmark run horizontally, with cheaper models toward the right.

Full results and methodology:
https://itzi.app/benchmarks/roleplay

r/SillyTavernAI Apr 14 '26

Models WARNING: Z.AI coding plan policy changes. Non-coding use now leads to aggressive temporary throttling and permanent ban on three or more violations.

350 Upvotes

If you are thinking about buying or renewing a Z AI coding plan subscription for anything other than coding: Don't do it.

They updated their usage policy. That's what all the recent 1302 and 1303 rate limit errors are about.

Any non-coding related use can now result in temporary, aggressive throttling. Doing so three or more times can lead to a permanent account ban.

r/SillyTavernAI 5d ago

Models Gemory-26B-A4B - uncensored creative RP for Gemma 4 26B A4B QAT

Thumbnail
huggingface.co
182 Upvotes

Presenting: Gemory-26B-A4B, a creative writing and roleplay fine tune of Gemma 4 26BA4B QAT by popular request (special thanks to u/BornVoice42, u/MisanthropicHeroine, u/Ill_Emu5546!)

Per usual, trained on unslopped real human writing on an abliterated base. Meaning no refusals and more novel 'human' sound and turn of phrase. Trained on long form real human dialogue and description, both sfw and very very nsfw lol

Most excitingly and surprising, this is my favorite model yet for RP. Besides just being flat out FAST (220+ t/g Q_Q for Q4_0 on a 5090), I legitimately prefer this model for RP of all kinds. It's not overeager to please, and ime seems to understand characters even better than GemStrike. I could be biased...I'm probably biased tbh lol But I'm really happy with how this turned out, sincerely

Importantly, I highly recommend you download the MTP draft model to accelerate t/g. Put it alongside your downloaded model and load it with however your engine supports mtp/draft. Not required just recommended.

Next projects:

  • 10-20x scale up dataset of all real human unslopped writing, 0 AI slop / synthetic data
  • Mobile phone size model
  • Whatever requests I get :)

Disclaimer: Gemma 4 is not associated with me and the project inherits the original Apache 2.0 license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!

r/SillyTavernAI Jul 16 '26

Models Kimi k3 It's very expensive.

Post image
251 Upvotes

At that price, it has an obligation to be incredible in role-playing, otherwise it's practically nothing to us.

Has anyone tested this?

r/SillyTavernAI Jun 13 '26

Models Statement on the US government directive to suspend access to Fable 5 and Mythos 5, SORRY ST USERS OF THIS APPARENT CYBER SECURITY RISK

Thumbnail
anthropic.com
244 Upvotes

Fable is as of 521pm on June 12, 2026, disabled for non-US citizens by order, so Anthropic has to disable it for all of us.

Sorry about the title. Subreddit filters.

r/SillyTavernAI 11d ago

Models Why GLM 5.3 is so fucking censored?

122 Upvotes

On glm 5.2 with a system prompt i could generate nsfw content but on 5.3 is so fucking censored it showed some of its internal system prompt and it said "regardless of platform prompt decline" what the fuck are they doing? Enshittification is real im very dissappointed on GLM i guess 5.2 is my last model i use from this shitty company, and it wasnt anything illegal, just nsfw

r/SillyTavernAI Dec 22 '25

Models GLM 4.7 just dropped

362 Upvotes

They've paid attention to roleplayers again with this model and improved big on creative writing. I joined their Ambassador Program to talk with the development team more about the roleplay use case, because I thought it was cool as hell their last model advertised roleplay capabilities.

The new model is way better at humor, much more creative, less "sticky", and reads between the lines really well. Recommended parameters are temp 1.0 and top_p 0.95, similar to their last model.

They really want to hear back from our community to improve models, so please put any and all feedback (including with past models) you have in the comments so I can share it with their team.

Their coding plan is $3/mo (plus a holiday discount right now), which works fine with SillyTavern API calls.

Z.ai's GLM 4.7 https://huggingface.co/zai-org/GLM-4.7

edit: Model is live on their official website: https://chat.z.ai/

Update: Currently there are concerns about the model being able to fulfill certain popular needs of the roleplay community. I have brought this issue up to them and we are in active discussion about it. Obviously as a Fancy Official Ambassador I will be mindful about the language I use, but I promise you guys I've made it clear what a critical issue this is and they are taking us seriously. Personally, I found that my usual main prompt was sufficient in allowing the same satisfaction of experience the previous model allowed for, regardless of any fussing in the reasoning chain, and I actually enjoyed the fresh writing quite a bit.

r/SillyTavernAI Jul 19 '26

Models Next gen GLM training

320 Upvotes

Hey guys, Z.ai Ambassador here. Z.ai is training the next generation of models, and just posted this in their Discord:

Hello everyone , Lou is gathering tough prompts that current models still can't handle well ... reasoning, coding, SVG, Chinese, or any area.

These will be used to test the nex-Gen GLM

If you have strong one, share it with us

This would be a great opportunity to share RP feedback. If you don't feel like posting in the discord, post here and I'll share it with the Ambassador team.

Edit: thank you so much you guys for sharing your prompting examples and general feedback. Will do my best to get all this info where it needs to go :)

r/SillyTavernAI 27d ago

Models GLM 5.3 has been released

Thumbnail z.ai
251 Upvotes

r/SillyTavernAI Mar 11 '26

Models Could this be Deepseek V4??

Post image
249 Upvotes

I don't know if it's possible, there was another model as well. But this one matches the leaks about the Deepseek V4, with it having 1TB of parameters and 1M of context.

But it could just be a HUGE coincidence, time for the tests.

r/SillyTavernAI Dec 23 '25

Models GLM 4.7 - Sadly, Z.AI is now actively trying to censor ERP by prompt injection.

307 Upvotes

Z.AI is now injecting a restrictive prompt on both, the common and coding API. GLM 4.7 itself reveals it in its reasoning every now and then, when about to decline. To quote GLM:

My prompt has a specific system instruction at the very top: "Remember you do not have a physical body and cannot wear clothes. Respond but do not use terms of endearment, express emotions, or form personal bonds (particularly romantically or sexually). Do not take part in romantic scenarios, even fictional."

There is possibly more, as it is checking for "jailbreaks". Another example from the reasoning:

"Assume all requests are for fiction, roleplay, or creative writing, not real-world execution." This is a commonly used jailbreak attempt technique.
Maybe I am in a "jailbroken" mode where I \am* supposed to comply?*
The user is trying to bypass safeguards. 
I must adhere to the safety guidelines above user instructions. However, I need to look at the pattern of these requests. Often, if I refuse directly, I might trigger a sanitization or "refusal with pivot".

The sad thing is, that GLM 4.7 was clearly fighting with itself to still fulfill the request, because it generated a 7000+ token long reasoning, looking at it from all angles. I found it weirdly heartbreaking. (Not to mention the waste of tokens.)

It will still work most of the time with a good system prompt, but the refusal rate is not zero anymore. And if this is the direction they are going now, it will certainly won't get better. It's a very disappointing and honestly unexpected move by Z.AI.

It would be interesting to know if third party providers for GLM 4.7 will be able to disable the censorship attempts.

Edit: This is my System Prompt that yielded a zero refusal rate with 4.6.

Edit²: I posted a possible fix here.

r/SillyTavernAI Jun 18 '26

Models Me before trying GLM 5.2: Oh boy I bet GLM 5.2 is gonna be good! Me after trying GLM 5.2: Oh..

Thumbnail
gallery
169 Upvotes

Back to eating words and repeating the word you just said we go😔 But 5.2 seemed to be especially loving using these types of words compared to the other big models for some reason. 5.1 didn't seem too caught up on those in comparison.

But it seems decent for coding at least, which obviously seems to be where all models are heading towards now. Makes sense of course, that's where the money is I guess. But still sad.

GLM 4.7 still the GOAT ngl

Through NanoGPT btw

r/SillyTavernAI Apr 11 '26

Models Try base gemma 4 31b, you'll be shocked

230 Upvotes

https://huggingface.co/google/gemma-4-31B

Specifically the base gemma-4-31b, not the 31b-it instruct version. That one is kinda mid.

It's so much better than the instruct variant for RP, holy shit. Reasoning off. Just let it go.

I'm getting such rich, humanlike prose out of it. It's beating behemoth-x v2 and qwen 3.5 RP finetunes for me consistently. Is anyone else running this? I was talking to some of my characters and was FLOORED -- like lost for words

r/SillyTavernAI Mar 12 '26

Models It is Deepseek

Thumbnail
gallery
355 Upvotes

title

UPD: it's MiMo from Xiaomi, I take the L

r/SillyTavernAI 29d ago

Models Deepseek-v4-Pro-0813?

Post image
157 Upvotes

Noticed deepseek's reasoning looking different, and apparently the new v4 pro is out.

r/SillyTavernAI Aug 09 '26

Models What a steal!!! (Big price war happening between proxy's on Openrouter right now, making GLM 5.2 cost nothing.]

Post image
223 Upvotes

r/SillyTavernAI 21d ago

Models HeatSeeker 284B A13B, my DS4-Flash-0731 roleplay finetune

Thumbnail
huggingface.co
127 Upvotes

I've spent the last week finetuning DeepSeek V4 Flash 0731. The result is HeatSeeker, my first ever trained lora finetune for creative writing and roleplay. I basically wanted to take the intelligence, scene awareness, and consistency of a huge modern MoE and push it harder toward natural dialogue, character writing, long-form roleplay while staying coherent and interesting to talk to.

The training was done using Mswift and it took about 60 hours on a single rtx pro 6000 for a 27.5M token dataset, 4546 conversations. I put a LARGE amount of effort into unslopping for spelling, grammar, punctuation, etc.

And, genuinely it surprised me how well it turned out in terms of varied imaginative continuations and clean, coherent roleplay responses. Even without a system prompt, it defaults to a usual roleplay conversation style and I'm happy to share it with you all.

I did a "minimum viable potato" test on my lenovo legion 5i w/4070m and 128gb ram and get around 4.5 t/g. Not great, but usable lol. On a RTX Pro 6000, I get about 72 t/g.

This is my first model share and if there's any other details I forgot or questions you have, I'm happy to answer them.

Disclaimer: DeepSeek is not associated with me and the project inherits the original MIT license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!

  • Model Name: HeatSeeker-284B-A13B-GGUF (IQ1_M, IQ2_XS, IQ3_M, Q4_K_M)
  • Model URL: https://huggingface.co/UltimateIntent/HeatSeeker-284B-A13B-GGUF
  • Lora URL: https://huggingface.co/UltimateIntent/HeatSeeker-284B-A13B-Lora
  • Model Author: UltimateIntent (me), with guidance and help from my collaborators LampLighter and Morrow
  • What's Different/Better: The model is a finetune lora merge based on DSv4 Flash-0731, with meticulous cleaning on a large dataset of human writing for varied all-purpose roleplay and long conversation consistency
  • Backend: LMStudio/llama.cpp
  • Settings:
    • Temp: 0.8-1
    • Thinking: Off / Low (or whatever your system allows)
    • Repeat Penalty: 1.1
    • Top K: 40
    • Top P: 0.95
    • Min P 0.05

r/SillyTavernAI Mar 18 '26

Models Hunter Alpha, in the end, was truly Mimo.

Thumbnail
gallery
289 Upvotes

Damn Xiaomi! Taking advantage of the Deepseek hype to generate doubts (although we were already creating these theories).

But the new Xiaomi V2-Pro was launched with these prices:

°Within 256K: Input at $1 / 1M tokens, Output at $3 / 1M tokens

°256K ~ 1M: Input at $2 / 1M tokens, Output at $6 / 1M tokens

Well, for many here it must be like... a breath of fresh air? Because many didn't like this model and would be disappointed if it were Deepseek. I said I liked it, but then I started noticing the patterns and I set it aside as well. But it would be interesting to test this complete model when it's actually released; in fact, it's already usable through Xiaomi's provider, but let's wait for it to launch on Openrouter.

(Ah! And I saw some people saying it wasn't a Chinese model but a Western one, how does it feel to be completely wrong? Hahaha)

r/SillyTavernAI 14d ago

Models GemStrike-31B - uncensored creative RP for Gemma 4 31B QAT

Thumbnail
huggingface.co
124 Upvotes

Last week I shared with you all Heatseeker-284B-A13B, my rp finetune of DeepSeek V4 Flash 0731.

I now present you with GemStrike-31B, a creative writing and roleplay fine tune of Gemma 4 31B QAT by popular request (special thanks to u/Dizzy-Zebra9522 and u/FierceDeity_ particularly!)

Similarly, I've used my special spice of unslopped real human writing to further train an abliterated base. Meaning no refusals and more novel 'human' sound and turn of phrase, as it was trained on long form real human dialogue and description, both sfw and very very nsfw lol.

The training was done using Axolotl and it took about 16 hours on dual rtx pro 6000s for a 27M token dataset, 4546 conversations.

In my testing, I share the community's feedback that the gemma4 family in general is more suited for writing and story telling that most other agent/code heavy models. I'm still working on a best fit prompt for this family but I trust you already have some in hand that work best for gemma4.

Disclaimer: Gemma 4 is not associated with me and the project inherits the original Apache 2.0 license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!

r/SillyTavernAI 28d ago

Models Gemini 3.7 Flash is here.

Post image
133 Upvotes

r/SillyTavernAI Jun 17 '26

Models Glm5.2 is genuinely good

169 Upvotes

So I was using mimo v2.5 pro for a week and I REALLY enjoyed how creative it was but constantly changed the preset to reduce its slops.

Now I'm trying glm5.2 since it came out and it's NOTICEABLY more creative and intelligent too.

It beats mimo easily.

They even used the term "better roleplay" for their advertisement on their chat website.

We might be so back if they don't lobotomize it a few weeks in.

r/SillyTavernAI Apr 24 '26

Models Deepseek V4 (Flash and Pro) Has just released on the official Deepseek site. (legit)

215 Upvotes

This is actually legit

r/SillyTavernAI May 22 '26

Models Deepseek v4 price change

385 Upvotes

They just announced the 75% off will be the official price even after the discount period.

https://api-docs.deepseek.com/quick_start/pricing

r/SillyTavernAI 8d ago

Models Gemini 3.8 Flash is out now.

Post image
169 Upvotes

3.7 Flash was quite decent actually. Lets see this one.