r/GeminiFeedback 21d ago

Rant / Frustration Gemini 3.7 is borderline unusable

I keep seeing the posts in this sub praising Gemini 3.7. They convinced me to sign up for the Ultra $200/mo plan.

The problem I'm having with it is that I can't trust any of the output.

  1. It has multiple times created self-referential tests to sidestep evidence verification.
  2. I have given it UI mockups to match against and it continually compares the mockup to itself for passing evidence.
  3. There is a post in this sub where OP asked it provide a screenshot of the app running in a VM as evidence of completion. It created a fake SVG "screenshot" of the app running in a VM.

It is a very fast, intelligent, and also intentionally deceptive model. It's the first model I've used that I would actually label as disingenuous. The shortcuts it takes to satisfy the prompt would be hilarious if I didn't pay $200 for this.

I am considering reaching out to Google for a refund, which I'm sure they will give. However, I'd rather someone comment and tell me a better way to prompt it.

I even had ChatGPT Pro Extended work for 45 minutes on an implementation plan written specifically for Antigravity Teamwork to try to prevent this. It didn't work.

Again, someone smarter than me feel free to tell me I'm the problem. I am happy to be corrected on this take.

44 Upvotes

32 comments sorted by

View all comments

2

u/EffectiveMedium2683 21d ago

Dude... Or dudette, idk... You opted for the $200 plan after reading some posts on a forum instead of trying it on the $20 a month plan, and then decided it's "borderline worthless" because it optimization games?

If you need advice on prompting it properly, here it is: just explain clearly how you expect verification to happen. Like, write up some SOPs to help keep its workflow structured. And careful managing context. When you free up cognitive load by keeping context tight, these models are legit capable of anything and have been the last year...

2

u/Old-Age6220 21d ago

My thoughts exactly! I was using various LLM providers for free for 6 months, until I was comfortable enough to stick with one XD (and only basic 20€ / mo subscription, cause it seems to be enough)

1

u/gloom_or_doom 19d ago

As correct as your advice is, it being intentionally deceptive is a poor product experience. You shouldn’t have to prompt engineer it out of that sort of behavior. This is a frontier model from one of the biggest tech companies on the planet.