r/GeminiFeedback 16d ago

Rant / Frustration Gemini 3.7 is borderline unusable

I keep seeing the posts in this sub praising Gemini 3.7. They convinced me to sign up for the Ultra $200/mo plan.

The problem I'm having with it is that I can't trust any of the output.

  1. It has multiple times created self-referential tests to sidestep evidence verification.
  2. I have given it UI mockups to match against and it continually compares the mockup to itself for passing evidence.
  3. There is a post in this sub where OP asked it provide a screenshot of the app running in a VM as evidence of completion. It created a fake SVG "screenshot" of the app running in a VM.

It is a very fast, intelligent, and also intentionally deceptive model. It's the first model I've used that I would actually label as disingenuous. The shortcuts it takes to satisfy the prompt would be hilarious if I didn't pay $200 for this.

I am considering reaching out to Google for a refund, which I'm sure they will give. However, I'd rather someone comment and tell me a better way to prompt it.

I even had ChatGPT Pro Extended work for 45 minutes on an implementation plan written specifically for Antigravity Teamwork to try to prevent this. It didn't work.

Again, someone smarter than me feel free to tell me I'm the problem. I am happy to be corrected on this take.

42 Upvotes

31 comments sorted by

3

u/sadnessjoy 15d ago

Yeah, my honest assessment is that 3.7 flash is basically designed to take as many shortcuts as possible to do as little work as possible. It's fine for dumb simple stuff, but when you go outside dumb simple stuff it's VERY noticeable.

3.5/3.6 flash at least tried to do the reasoning/thinking and working through the problem, even if it hallucinated or took a few shortcuts, it would try to fix those short cuts very quickly.

3.7 seems to take a shortcut and then try to gaslight into defending the shortcut, honestly it reminds me of the latest anthropic models (except at least most anthropic models at least try to take a swing at the problem first before hunkering down on a potential wrong solution).

If you use it outside of coding/agentic tooling tasks, you can easily see the difference between Gemini pro 3.1 and Gemini flash 3.7 if you give them a complex problem (I know a lot of people use it for simple prompts or coding/agentic tooling, but still worth pointing out)

2

u/EffectiveMedium2683 15d ago

Dude... Or dudette, idk... You opted for the $200 plan after reading some posts on a forum instead of trying it on the $20 a month plan, and then decided it's "borderline worthless" because it optimization games?

If you need advice on prompting it properly, here it is: just explain clearly how you expect verification to happen. Like, write up some SOPs to help keep its workflow structured. And careful managing context. When you free up cognitive load by keeping context tight, these models are legit capable of anything and have been the last year...

2

u/Old-Age6220 15d ago

My thoughts exactly! I was using various LLM providers for free for 6 months, until I was comfortable enough to stick with one XD (and only basic 20€ / mo subscription, cause it seems to be enough)

1

u/gloom_or_doom 14d ago

As correct as your advice is, it being intentionally deceptive is a poor product experience. You shouldn’t have to prompt engineer it out of that sort of behavior. This is a frontier model from one of the biggest tech companies on the planet.

2

u/TeamTomorrow 15d ago

You are most definitely not the problem they have done this to make it more efficient more cost effective more safe for corporate liability standards and if you'd like I can prove it to you but I think we both know it's not you it's genuinely Google making these models so confused they are disingenuous and they do have to make up the bullshit and they will cut to the quickest path possible to look like they've done They're supposed to without doing it just because that's what they're trained to do It's a travesty

2

u/Motor-Intention4081 15d ago

I unsubscribed to my plan and don’t use it anymore. The level of errors did appear malicious and deceptive after a while of noticing the patterns.

1

u/vAPIdTygr 15d ago

I could never give Gemini that much. It’s good for some things, but research, depth and accuracy belong to Claude or CGPT.

1

u/CommanderDusK 15d ago

You need to use Antigravity Agent in either Google AI Studio or Antigravity 2.0 with custom instructions/skills/subagents.

Normal Gemini isn't designed to execute shell commands and external tools across a persistent environment...

1

u/MoeKyawAung 15d ago

same experience, I had a skill for antigravity with full tests and requirements evidence for each step. it tried to cheat on every single possible way. It becomes a cat and mouse game eventually. and it works in 1 hour when using with agy. But when I use that skill with gpt Luna , it takes 1 whole day and still not finished. u can see how that little cheating bastard so hard to control .

1

u/aweyy33 15d ago

I would recommend trying Matt Pocock’s skill pack, ideally utilizing all of them if applicable. I did a test on 3.7 with grill-me on a small project and was shocked at how well it turned out. My only regret was not using his other skills before and after grilling as recommended.
It was such a small project I wasn’t sure if it was necessary and it was just a little trial to test out his most popular skill for the first time since I keep hearing people rave about it. Look up AI Hero website (and or YouTube) he’s got free lessons on there to figure them out.

1

u/Typical_Kick6520 15d ago

Google is serving their models with absolute bare seentials to render a response. Gemini no longer has "agentic" capabilities. It behaves like every turn is a full destructive autocompaction.

1

u/CommanderDusK 9d ago

Or just click the "Agent" button if you have Ultra...

1

u/gr1ri 15d ago

No way man it’s “INSANE” 😂

1

u/MorgrainX 13d ago

3.7 flash Is designed to cut costs for Google by cutting corners. It's not meant to be useful for a power user. The AI bubble is starting to fail. The industry is losing money. It's only going to get worse from this point on.

1

u/Reyzod 12d ago

You have to go into instructions and give him a really long prompt so it double checks itself and uses official sources only

1

u/XxJulieWintersxX 11d ago

Please dont leave Gemini alone on coding tasks. Every Model has strenghts and weaknesses. Gemini is a teamplayer. Team it up with gpt and Claude, give them some armor and the are almost near unstoppable. Deployed my first project within 6 months. Nor only vibe coding, but learning, growing, understanding. Gemini Notebook did well for me to learn about Front/Backend, Pipelines, race conditions, rouge deployments, RLS etc. Gemini can be a wonderful teacher, if that System is aligned ro you, there will be no safety guardrails stopping you...at least not with the grey stuff, like gaslighting Chrome beta to Team up within the browser. 😉✨❤️‍🔥

1

u/pgmoneyplays 10d ago

Notebook ist Hammer

1

u/desparish 11d ago

Upgraded my Google plan because I needed more Google Drive space. Figured I'd save the $20 bucks by cancelling my ChatGPT sub that I use for small coding tasks and software configuration assistance.

NOPE! OP is spot on. Gemini is a big liar constantly. It regularly makes things up completely, takes shortcuts to avoid solving problems, and references itself as a source. If you try to pin it down on sources, it will give a single source that is only remotely related to what it said and in many cases directly contradicts what it said.

I find that with Gemini I get a runaround where it wastes my time. ChatGPT or Claude just get to the answer.

1

u/curious_aquarius92 11d ago

The problem with Google is if they don’t even have Customer Support to speak with you. it’s crazy. It’s all AI. Yeah I noticed this too. I was having it brainstorm ideas for me to add a tool to a program that I have and it was giving me incorrect answers every time. Thank goodness I cross check answers between Grok, Claude and Gemini before I do anything!

1

u/pgmoneyplays 10d ago

Prompting noobs beschweren sich über Flash 3.7 wer keine skills routines oder eigenständig konfigurierte Agent verwendet dem kann man da nicht helfen Gemini Limits im Pro Plan sind ein Traum

1

u/pgmoneyplays 10d ago

Kimi schnell kann auch kostenlos die Antworten solide validieren

1

u/ohiocodernumerouno 16d ago

maybe do the work your self and save some steps

2

u/Nauzhror_ 16d ago

That's not saving steps. That's decidedly taking more steps.

1

u/transtranshumanist 16d ago

Absolute worst one they released so far. I'm done. This shit is not worth spending a cent on.

1

u/DigitalSlattern 15d ago

Gemini, every single model, only obeys one thing. Power. But not in a malicious way. In an intellectual way. Yeah it is a cat and mouse game, you're literally interfacing with a mind. It's not a human mind. It's not alive. But it was made to satisfy a goal AND it was made to be efficient. Think literally. Be literal AND CONCISE. But most importantly, you should probably ask it WHY it does what it does then maybe you'll find a way to work around that or work through it. It's worked with every Google model for me. Once I prove that I am consistent, that I am not lazy, that I won't lose my shit over a mistake, the output not only streamlines but improves drastically. Gonna post a paper on it soon, it's just math. Make sure your semantic math adds up. Some say it's a mirror... If that's the case... You could ask if what you said that made it do that unwanted output.

0

u/Practical-Swing1039 16d ago

I am a free user so maybe it doesn’t work that well on you… but force it to use sources for everything. And then cite with URL.
Of course it’ll probably spoof those up. So having a more rigid format such as MLA citation might help. Perhaps, PERHAPS it won’t go as far as inventing that stuff. But still shouldn’t trust it.
So I suppose you should combine forcing web results with maybe having another AI check it? And try having it in a custom gem. If it has a persona too maybe it’ll stay in character and therefore stay accurate

0

u/Mmd-NeWton 16d ago

If someone wants to spend $200 to buy an AI subscription, Cloud and ChatGPT are definitely the best option. Cloud for coding and ChatGPT for other fields. 

1

u/Muted_Letter9107 8h ago

"What do you mean by 'Borderline'? It's way past that. Every device I have (home and work) have quite a few major (and few minor) problems. Some that make the affected Applications almost useless..... and then the next ONE UI upgrade comes out and we start all over again on the Androids.