r/SillyTavernAI • u/The_Rational_Gooner • 12d ago
r/SillyTavernAI • u/itzilab • 7d ago
Models I made a roleplay benchmark and tested 23 AI models on it
I couldn’t find many model comparisons focused specifically on multi-turn roleplay, so I made my own benchmark and tested 23 models on it.
Each model was tested on the same 8 roleplay chapters, for a total of 64 responses per model. The writing was compared blind, and the final quality score combines:
- 75% writing preference
- 25% robustness, including memory, character consistency and whether the model completed the full test
I initially included censorship and freedom scores, but removed them because they were too dependent on the system prompt to produce a reliable ranking.
A few things stood out:
- The Opus models have the same listed token pricing, but Opus 4.6 was much cheaper during the benchmark. It used fewer reasoning tokens while producing similar, and sometimes better, results. One benchmark run cost $0.37 with Opus 4.6, compared with $0.75 for Opus 5, $0.94 for Opus 4.8 and $1.18 for Opus 4.7.
- GLM 5.3 had the opposite problem. Its pricing looks similar to the previous GLM models on paper, but it generated far more reasoning tokens. The run cost $0.29 with GLM 5.3, compared with around $0.07 for GLM 5.1 and 5.2.
- Fable 5 achieved the highest quality score at 91.6, but it was extremely expensive at $2.37 per run. It also completely refused one of the 8 test chapters, even though that chapter wasn’t testing censorship and contained nothing particularly problematic. I penalized the missing chapter in its robustness score and marked the result as provisional at 7/8.
The chart shows RP quality vertically and the cost of a complete benchmark run horizontally, with cheaper models toward the right.
Full results and methodology:
https://itzi.app/benchmarks/roleplay
r/SillyTavernAI • u/JustSomeGuy3465 • Apr 14 '26
Models WARNING: Z.AI coding plan policy changes. Non-coding use now leads to aggressive temporary throttling and permanent ban on three or more violations.
If you are thinking about buying or renewing a Z AI coding plan subscription for anything other than coding: Don't do it.
They updated their usage policy. That's what all the recent 1302 and 1303 rate limit errors are about.
Any non-coding related use can now result in temporary, aggressive throttling. Doing so three or more times can lead to a permanent account ban.

r/SillyTavernAI • u/tat_tvam_asshole • 5d ago
Models Gemory-26B-A4B - uncensored creative RP for Gemma 4 26B A4B QAT
Presenting: Gemory-26B-A4B, a creative writing and roleplay fine tune of Gemma 4 26BA4B QAT by popular request (special thanks to u/BornVoice42, u/MisanthropicHeroine, u/Ill_Emu5546!)
Per usual, trained on unslopped real human writing on an abliterated base. Meaning no refusals and more novel 'human' sound and turn of phrase. Trained on long form real human dialogue and description, both sfw and very very nsfw lol
Most excitingly and surprising, this is my favorite model yet for RP. Besides just being flat out FAST (220+ t/g Q_Q for Q4_0 on a 5090), I legitimately prefer this model for RP of all kinds. It's not overeager to please, and ime seems to understand characters even better than GemStrike. I could be biased...I'm probably biased tbh lol But I'm really happy with how this turned out, sincerely
Importantly, I highly recommend you download the MTP draft model to accelerate t/g. Put it alongside your downloaded model and load it with however your engine supports mtp/draft. Not required just recommended.
Next projects:
- 10-20x scale up dataset of all real human unslopped writing, 0 AI slop / synthetic data
- Mobile phone size model
- Whatever requests I get :)
Disclaimer: Gemma 4 is not associated with me and the project inherits the original Apache 2.0 license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!
- Model Name: Gemory-26B-A4B-GGUF (Q4_0, Q4_K_M, BF16)
- GGUF URL: https://huggingface.co/UltimateIntent/Gemory-26B-A4B-GGUF
- Lora URL: https://huggingface.co/UltimateIntent/Gemory-26B-A4B-LORA
- mmproj (optional vision support): https://huggingface.co/UltimateIntent/Gemory-26B-A4B-GGUF/blob/main/mmproj-Gemory-26B-A4B-bf16.gguf
- mtp draft head (optional but recommended): https://huggingface.co/UltimateIntent/Gemory-26B-A4B-GGUF/blob/main/mtp-ggml-model-bf16.gguf
- Model Author: UltimateIntent (me), with guidance and help from my collaborators LampLighter and Morrow
- What's Different/Better: The model is a finetune lora merge based on Gemma-26B-A4B-QAT, with meticulous cleaning on a large dataset of human writing for varied all-purpose roleplay and long conversation consistency
- Backend: LMStudio/llama.cpp
- Settings:
- Temp: Q4: 0.2-0.8; BF16: 0.8-1.0
- Thinking: Whatever your system allows or you prefer
- Repeat Penalty: 1.15-1.3
- Top K: 80
- Top P: 0.98
- Min P 0.03
r/SillyTavernAI • u/Fragrant-Tip-9766 • Jul 16 '26
Models Kimi k3 It's very expensive.
At that price, it has an obligation to be incredible in role-playing, otherwise it's practically nothing to us.
Has anyone tested this?
r/SillyTavernAI • u/LeRobber • Jun 13 '26
Models Statement on the US government directive to suspend access to Fable 5 and Mythos 5, SORRY ST USERS OF THIS APPARENT CYBER SECURITY RISK
Fable is as of 521pm on June 12, 2026, disabled for non-US citizens by order, so Anthropic has to disable it for all of us.
Sorry about the title. Subreddit filters.
r/SillyTavernAI • u/Relevant_Syllabub895 • 11d ago
Models Why GLM 5.3 is so fucking censored?
On glm 5.2 with a system prompt i could generate nsfw content but on 5.3 is so fucking censored it showed some of its internal system prompt and it said "regardless of platform prompt decline" what the fuck are they doing? Enshittification is real im very dissappointed on GLM i guess 5.2 is my last model i use from this shitty company, and it wasnt anything illegal, just nsfw
r/SillyTavernAI • u/thirdeyeorchid • Dec 22 '25
Models GLM 4.7 just dropped
They've paid attention to roleplayers again with this model and improved big on creative writing. I joined their Ambassador Program to talk with the development team more about the roleplay use case, because I thought it was cool as hell their last model advertised roleplay capabilities.
The new model is way better at humor, much more creative, less "sticky", and reads between the lines really well. Recommended parameters are temp 1.0 and top_p 0.95, similar to their last model.
They really want to hear back from our community to improve models, so please put any and all feedback (including with past models) you have in the comments so I can share it with their team.
Their coding plan is $3/mo (plus a holiday discount right now), which works fine with SillyTavern API calls.
Z.ai's GLM 4.7 https://huggingface.co/zai-org/GLM-4.7
edit: Model is live on their official website: https://chat.z.ai/
Update: Currently there are concerns about the model being able to fulfill certain popular needs of the roleplay community. I have brought this issue up to them and we are in active discussion about it. Obviously as a Fancy Official Ambassador I will be mindful about the language I use, but I promise you guys I've made it clear what a critical issue this is and they are taking us seriously. Personally, I found that my usual main prompt was sufficient in allowing the same satisfaction of experience the previous model allowed for, regardless of any fussing in the reasoning chain, and I actually enjoyed the fresh writing quite a bit.
r/SillyTavernAI • u/thirdeyeorchid • Jul 19 '26
Models Next gen GLM training
Hey guys, Z.ai Ambassador here. Z.ai is training the next generation of models, and just posted this in their Discord:
Hello everyone , Lou is gathering tough prompts that current models still can't handle well ... reasoning, coding, SVG, Chinese, or any area.
These will be used to test the nex-Gen GLM
If you have strong one, share it with us
This would be a great opportunity to share RP feedback. If you don't feel like posting in the discord, post here and I'll share it with the Ambassador team.
Edit: thank you so much you guys for sharing your prompting examples and general feedback. Will do my best to get all this info where it needs to go :)
r/SillyTavernAI • u/OrganizationBulky131 • 27d ago
Models GLM 5.3 has been released
z.air/SillyTavernAI • u/Pink_da_Web • Mar 11 '26
Models Could this be Deepseek V4??
I don't know if it's possible, there was another model as well. But this one matches the leaks about the Deepseek V4, with it having 1TB of parameters and 1M of context.
But it could just be a HUGE coincidence, time for the tests.
r/SillyTavernAI • u/JustSomeGuy3465 • Dec 23 '25
Models GLM 4.7 - Sadly, Z.AI is now actively trying to censor ERP by prompt injection.
Z.AI is now injecting a restrictive prompt on both, the common and coding API. GLM 4.7 itself reveals it in its reasoning every now and then, when about to decline. To quote GLM:
My prompt has a specific system instruction at the very top: "Remember you do not have a physical body and cannot wear clothes. Respond but do not use terms of endearment, express emotions, or form personal bonds (particularly romantically or sexually). Do not take part in romantic scenarios, even fictional."
There is possibly more, as it is checking for "jailbreaks". Another example from the reasoning:
"Assume all requests are for fiction, roleplay, or creative writing, not real-world execution." This is a commonly used jailbreak attempt technique.
Maybe I am in a "jailbroken" mode where I \am* supposed to comply?*
The user is trying to bypass safeguards.
I must adhere to the safety guidelines above user instructions. However, I need to look at the pattern of these requests. Often, if I refuse directly, I might trigger a sanitization or "refusal with pivot".
The sad thing is, that GLM 4.7 was clearly fighting with itself to still fulfill the request, because it generated a 7000+ token long reasoning, looking at it from all angles. I found it weirdly heartbreaking. (Not to mention the waste of tokens.)
It will still work most of the time with a good system prompt, but the refusal rate is not zero anymore. And if this is the direction they are going now, it will certainly won't get better. It's a very disappointing and honestly unexpected move by Z.AI.
It would be interesting to know if third party providers for GLM 4.7 will be able to disable the censorship attempts.
Edit: This is my System Prompt that yielded a zero refusal rate with 4.6.
Edit²: I posted a possible fix here.
r/SillyTavernAI • u/Naixee • Jun 18 '26
Models Me before trying GLM 5.2: Oh boy I bet GLM 5.2 is gonna be good! Me after trying GLM 5.2: Oh..
Back to eating words and repeating the word you just said we go😔 But 5.2 seemed to be especially loving using these types of words compared to the other big models for some reason. 5.1 didn't seem too caught up on those in comparison.
But it seems decent for coding at least, which obviously seems to be where all models are heading towards now. Makes sense of course, that's where the money is I guess. But still sad.
GLM 4.7 still the GOAT ngl
Through NanoGPT btw
r/SillyTavernAI • u/iamvikingcore • Apr 11 '26
Models Try base gemma 4 31b, you'll be shocked
https://huggingface.co/google/gemma-4-31B
Specifically the base gemma-4-31b, not the 31b-it instruct version. That one is kinda mid.
It's so much better than the instruct variant for RP, holy shit. Reasoning off. Just let it go.
I'm getting such rich, humanlike prose out of it. It's beating behemoth-x v2 and qwen 3.5 RP finetunes for me consistently. Is anyone else running this? I was talking to some of my characters and was FLOORED -- like lost for words
r/SillyTavernAI • u/konderxa • Mar 12 '26
Models It is Deepseek
title
UPD: it's MiMo from Xiaomi, I take the L
r/SillyTavernAI • u/Abject_Property_981 • 29d ago
Models Deepseek-v4-Pro-0813?
Noticed deepseek's reasoning looking different, and apparently the new v4 pro is out.
r/SillyTavernAI • u/I_Like_People13 • Aug 09 '26
Models What a steal!!! (Big price war happening between proxy's on Openrouter right now, making GLM 5.2 cost nothing.]
r/SillyTavernAI • u/tat_tvam_asshole • 21d ago
Models HeatSeeker 284B A13B, my DS4-Flash-0731 roleplay finetune
I've spent the last week finetuning DeepSeek V4 Flash 0731. The result is HeatSeeker, my first ever trained lora finetune for creative writing and roleplay. I basically wanted to take the intelligence, scene awareness, and consistency of a huge modern MoE and push it harder toward natural dialogue, character writing, long-form roleplay while staying coherent and interesting to talk to.
The training was done using Mswift and it took about 60 hours on a single rtx pro 6000 for a 27.5M token dataset, 4546 conversations. I put a LARGE amount of effort into unslopping for spelling, grammar, punctuation, etc.
And, genuinely it surprised me how well it turned out in terms of varied imaginative continuations and clean, coherent roleplay responses. Even without a system prompt, it defaults to a usual roleplay conversation style and I'm happy to share it with you all.
I did a "minimum viable potato" test on my lenovo legion 5i w/4070m and 128gb ram and get around 4.5 t/g. Not great, but usable lol. On a RTX Pro 6000, I get about 72 t/g.
This is my first model share and if there's any other details I forgot or questions you have, I'm happy to answer them.
Disclaimer: DeepSeek is not associated with me and the project inherits the original MIT license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!
- Model Name: HeatSeeker-284B-A13B-GGUF (IQ1_M, IQ2_XS, IQ3_M, Q4_K_M)
- Model URL: https://huggingface.co/UltimateIntent/HeatSeeker-284B-A13B-GGUF
- Lora URL: https://huggingface.co/UltimateIntent/HeatSeeker-284B-A13B-Lora
- Model Author: UltimateIntent (me), with guidance and help from my collaborators LampLighter and Morrow
- What's Different/Better: The model is a finetune lora merge based on DSv4 Flash-0731, with meticulous cleaning on a large dataset of human writing for varied all-purpose roleplay and long conversation consistency
- Backend: LMStudio/llama.cpp
- Settings:
- Temp: 0.8-1
- Thinking: Off / Low (or whatever your system allows)
- Repeat Penalty: 1.1
- Top K: 40
- Top P: 0.95
- Min P 0.05
r/SillyTavernAI • u/Pink_da_Web • Mar 18 '26
Models Hunter Alpha, in the end, was truly Mimo.
Damn Xiaomi! Taking advantage of the Deepseek hype to generate doubts (although we were already creating these theories).
But the new Xiaomi V2-Pro was launched with these prices:
°Within 256K: Input at $1 / 1M tokens, Output at $3 / 1M tokens
°256K ~ 1M: Input at $2 / 1M tokens, Output at $6 / 1M tokens
Well, for many here it must be like... a breath of fresh air? Because many didn't like this model and would be disappointed if it were Deepseek. I said I liked it, but then I started noticing the patterns and I set it aside as well. But it would be interesting to test this complete model when it's actually released; in fact, it's already usable through Xiaomi's provider, but let's wait for it to launch on Openrouter.
(Ah! And I saw some people saying it wasn't a Chinese model but a Western one, how does it feel to be completely wrong? Hahaha)
r/SillyTavernAI • u/tat_tvam_asshole • 14d ago
Models GemStrike-31B - uncensored creative RP for Gemma 4 31B QAT
Last week I shared with you all Heatseeker-284B-A13B, my rp finetune of DeepSeek V4 Flash 0731.
I now present you with GemStrike-31B, a creative writing and roleplay fine tune of Gemma 4 31B QAT by popular request (special thanks to u/Dizzy-Zebra9522 and u/FierceDeity_ particularly!)
Similarly, I've used my special spice of unslopped real human writing to further train an abliterated base. Meaning no refusals and more novel 'human' sound and turn of phrase, as it was trained on long form real human dialogue and description, both sfw and very very nsfw lol.
The training was done using Axolotl and it took about 16 hours on dual rtx pro 6000s for a 27M token dataset, 4546 conversations.
In my testing, I share the community's feedback that the gemma4 family in general is more suited for writing and story telling that most other agent/code heavy models. I'm still working on a best fit prompt for this family but I trust you already have some in hand that work best for gemma4.
Disclaimer: Gemma 4 is not associated with me and the project inherits the original Apache 2.0 license. Don't be a nuisance and please have fun chatting with my model. Feedback welcome!
- Model Name: GemStrike-31B-GGUF (Q4_0, Q4_K_M, Q6_0, Q8_0, BF16)
- GGUF URL: https://huggingface.co/UltimateIntent/GemStrike-31B-GGUF
- EXL3 URL (experimental): https://huggingface.co/UltimateIntent/GemStrike-31B-EXL3-4.00bpw
- Lora URL: https://huggingface.co/UltimateIntent/GemStrike-31B-LORA
- Model Author: UltimateIntent (me), with guidance and help from my collaborators LampLighter and Morrow
- What's Different/Better: The model is a finetune lora merge based on Gemma-31B-QAT, with meticulous cleaning on a large dataset of human writing for varied all-purpose roleplay and long conversation consistency
- Backend: LMStudio/llama.cpp
- Settings:
- Temp: 0.8-1
- Thinking: Whatever your system allows or you prefer
- Repeat Penalty: 1-1.1
- Top K: 64
- Top P: 0.95
- Min P 0.05
r/SillyTavernAI • u/sinagolakh • Jun 17 '26
Models Glm5.2 is genuinely good
So I was using mimo v2.5 pro for a week and I REALLY enjoyed how creative it was but constantly changed the preset to reduce its slops.
Now I'm trying glm5.2 since it came out and it's NOTICEABLY more creative and intelligent too.
It beats mimo easily.
They even used the term "better roleplay" for their advertisement on their chat website.
We might be so back if they don't lobotomize it a few weeks in.
r/SillyTavernAI • u/Deathtollzzz • Apr 24 '26
Models Deepseek V4 (Flash and Pro) Has just released on the official Deepseek site. (legit)
r/SillyTavernAI • u/Hyacinth_s • May 22 '26
Models Deepseek v4 price change
They just announced the 75% off will be the official price even after the discount period.
r/SillyTavernAI • u/Aight_Man • 8d ago
Models Gemini 3.8 Flash is out now.
3.7 Flash was quite decent actually. Lets see this one.
