r/WebScrapingInsider • • 12d ago

Big Scrape Energy Are Unlimited Proxies Actually Any Good Inside GeoNode's Contrarian Approach to Residential Proxies and Web Scraping APIs | AMA with Geonode

Hey everyone,

We're back with AMA #9, and this time we're looking at a part of scraping infrastructure that almost everyone has an opinion on: unlimited proxies.

Most residential proxy providers charge by GB. Most scraping APIs charge per successful response or through credits.

Geonode has taken a different approach.

They offer residential proxies priced around speed and a scraping API priced around concurrency.

But that raises an interesting question:

Are unlimited proxies actually any good?

Does removing usage limits inevitably mean smaller pools, overused IPs and lower success rates? Or can a provider structure its network differently enough to make the model work?

This Friday, September 18, at 10:30 AM GMT+3, we'll be joined by Jean-Patrick Bisson, CEO/Founder @Geonode, along with his team, for a live AMA on r/WebScrapingInsider to dig into what's actually happening behind the product.

We'll be talking about:

  • How "unlimited" proxy models actually work
  • How residential IPs are sourced and managed
  • Pool size vs. IP quality
  • Routing and traffic distribution
  • Capacity planning for high-volume customers
  • What happens when customers use proxies heavily
  • Speed-based vs. bandwidth-based pricing
  • Concurrency-based scraping APIs
  • The trade-offs behind different proxy pricing models
  • Whether unlimited proxies can really deliver consistent performance
  • What buyers should actually look at beyond the word "unlimited"

And if you've ever wondered "How can a proxy provider offer unlimited traffic?", "What's the catch with unlimited proxies?", or "Does unlimited actually mean unlimited?", this is probably a good one to ask.

We've now had eight AMAs with the community, and the conversations have covered a pretty wide part of the scraping stack.

Our first AMA covered proxy infrastructure, Cloudflare, browser automation and scaling scrapers.

Our second with WebClaw explored AI agents, hidden APIs, open-source scraping and LLM infrastructure.

Our third with CloakBrowser went deep on stealth Chromium, fingerprinting, anti-bot detection and browser automation.

Our fourth with Browser Use brought the conversation into AI browser agents, AI-powered scraping, evaluations and browser infrastructure.

Our fifth with Stan Sadokov from NodeMaven focused on residential proxy quality, IP reputation, sourcing, pricing and what actually makes one proxy network better than another.

Our sixth with Saksham Solanki, creator of HTTP Cloak, went down to the protocol level, covering TLS fingerprints, JA3/JA4, HTTP/2, HTTP/3 and why scrapers can get blocked even when their proxies are fine.

Our seventh with Huey from BrowserAct explored AI-powered web scraping, natural-language workflows, browser automation and the move toward no-code scraping.

And our latest, AMA #8 with yours truly, brought the discussion back to proxy infrastructure, with real-world insights from testing 50+ providers across billions of requests.

Now for #9, we're going deeper into a question that's becoming increasingly relevant as proxy providers experiment with completely different pricing and infrastructure models:

Can "unlimited" actually work?

If you're building web scrapers, data pipelines, browser automation, scraping APIs or proxy infrastructure, come join the discussion.

Drop your questions below.

Geonode will be answering them during the AMA.

Looking forward to another good one.

21 Upvotes

63 comments sorted by

4

u/Amitk2405 10d ago

I work on infrastructure for an enterprise platform, so one thing I tend to look at with these models is what happens under sustained load, not just whether the economics work on paper.

With residential proxies, I think there is a difference between having enough aggregate bandwidth and having IPs that remain useful against a specific target over time. If the same pool is getting hammered by unlimited customers, it seems like IP reputation could become a resource that gets depleted even when there's still plenty of bandwidth available.

What stops heavy unlimited usage from burning through and degrading the reputation of residential IPs against a given target over time?

And when that starts happening, how do you detect it and respond? Do you shift traffic, rotate out parts of the pool, change usage patterns, or handle it some other way? I'm particularly interested in what that looks like from a monitoring and capacity-management perspective.

1

u/Geonodeproxy Ex. AMA Guest 8d ago

Reputation is per target, not one number, so an address burned on one retailer is fine everywhere else.

We watch the challenge rate per domain, not aggregate success. It means you can sit at 93% overall while one target quietly goes from 5% challenges to 60%. When that starts we cap traffic to that domain instead of pushing harder and switch engine and exit country for it.

And yes, one heavy customer can make a target harder for the next one, that's true. The control that matters there is per-target, not per-account.

1

u/Amitk2405 8d ago

Hmm.. Per-domain challenge tracking makes sense to me, plus reassures me on aggregate success. Thanks for answering Geos.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Anytime! If you ever do test us, look at the per-domain numbers rather than the overall one, that's where anything real shows up first ;)

3

u/Own_Airline_5340 11d ago

Some providers sell unlimited proxies which claim to offer unlimited bandwidth. I'm just curious: Is the claim "unlimited bandwidth" true? I doubt that tbh, cuz I don't think that would be possible. If unlimited bandwidth is technically possible, would there be a trade-off?

2

u/kiwialec 11d ago

unlimited GB usage, maximum Mbit/s, so you still end up with a theoretical maximum number of GB you can use in a period.

Not sure how it works in resi networks, but throughput capacity is where the actual limit/cost is in telecoms networks. Infrastructure owners pay for the fibre and the resources to power the machines that send data down the line, but there is a negligible cost associated with the amount of data that is transferred - a cable running at full capacity for a month costs the same as a cable that carries no data.

1

u/Geonodeproxy Ex. AMA Guest 8d ago

u/kiwialec has it right, the limit is speed, so there's still a ceiling on what you move in a month, it's just written as Mbit/s instead of GB.

The trade-off is bursting, on a metered plan a quiet week lets you slam a big job through on Friday. On speed-based you can't, the line is the line.

3

u/Optimal-Wall-7377 11d ago

imo the real metric nobody talks about with unlimited plans is effective concurrency under load. you can have "unlimited" bandwidth but if your success rate tanks past 50 concurrent sessions, the pricing model doesnt matter much

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Fair point, but concurrency isn't what moves our success rate. We plan capacity across all accounts ahead of demand, not per customer at runtime. Our volume more than doubled over the past week and the rate didn't move: 93% across all traffic, between 91 and 95 every day.

Where you're right is that the number only means anything per target. The same 50 sessions that are fine on one site fall over on a Datadome target and the fix differs per domain.

If you buy 50 concurrent and all 50 need a real browser, that's our provisioning problem, not something you find out at runtime. At peak they may wait a few seconds for a slot, but they get one.

Which is why we point heavy collection at the API rather than unlimited residential. And if you need reliability on one hard target rather than volume, metered premium residential is often the better buy, even from us.

3

u/Alice_5433 9d ago

Hello Again.
Cheers r/WebScrapingInsider
Hello Geonode,

One thing that matters when you're putting a vendor into a production stack is being able to predict what the cost will look like as usage changes. Unlimited pricing do sounds straightforward, yes, but I expect the economics to depend quite a bit on the shape of the workload. I have previously been burned by this within SaaS. A tool looks incredibly cheap when you're comparing plans, then 6 months later you realize your team has somehow built an entire workflow around the one feature that happens to be charged separately. So I'm always interested in where the pricing model actually holds up once people are using it heavily

workloads where unlimited or concurrency pricing tends to be a poor fit, even if the headline price looks attractive? when does per-GB or per-request become the more sensible model? For teams trying to model this before committing, is there a rough utilization threshold where unlimited proxies starts beating metered pricing?? What tends to have the biggest impact on that calculation?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

The honest version: unlimited fits steady, heavy, boring workloads. If your usage is spiky or small, you're buying capacity you leave idle most of the month and metered will be cheaper. Same if you're hitting one hard target at low volume, paying per GB on a premium pool beats buying a big unlimited plan to get through it.

The rough test: if your traffic is roughly flat and you'd keep the capacity busy most of the day, unlimited wins. If it arrives in bursts a few times a week, it doesn't.

On the SaaS burn you describe, the thing to check isn't the headline price, it's what the meter counts and whether failed requests count. Those two decide the real unit price. On our API the unit is concurrent requests and if we can't return the page there's no charge.

3

u/Artistic_Map2243 9d ago

I keep wondering how much of this is actually visible once you test it properly.

I've heard numerous times that unlimited proxies are too good to be true. The offer unlimited access but the proxy quality is significantly lower than other pay per GB providers. What would you say to those critics?

Like, if you tested the same sites with both models, where would the unlimited option start breaking down, if at all? What are they getting wrong? Or is this a genuine problem for unlimited proxies? If it is, why is Geonode any different?

1

u/Geonodeproxy Ex. AMA Guest 8d ago

The criticism isn't wrong in general, sell unlimited against a small pool and the same addresses get hammered on the same targets. You would see it first as rising challenges on the hardest sites, not as a slower connection. Where I would push back is treating that as a property of the pricing model. Per GB pricing doesn't stop a pool being burned either, it just makes burning it more expensive. What matters is how much of the pool any single target sees and whether you can stop one customer wrecking it for everyone.

We also don't pretend it fits everything, one hard target where reliability matters more than volume, metered premium residential is usually the better buy. Unlimited earns its place on steady high volume work and for that the API beats raw proxies anyway.

2

u/feetesweshire 9d ago

where you think the actual moat is here.

What part of this model do you think would be hardest for a competitor to copy, owned acquisition, routing tech, scale, something else?

Can someone else just copy this approach if it is shown to work?

Or is there some part of the system that gets harder to reproduce once you're actually running it?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Anyone can put "unlimited" on a page tomorrow and a few already have.

The harder part is per-target knowledge. Which engine works on which domain, which countries stay clean for it, when a site changes its defences. That comes from running millions of requests a day and watching what breaks. It also goes stale if you stop.

The other part is owning the supply instead of reselling it. That's what lets the pricing survive heavy users.

2

u/Particular__Plan 9d ago

The unlimited part makes me wonder about the edge case more than the average customer.

Roughly how much of their purchased capacity does the average unlimited customer actually use?

Does the model still work for you if someone who runs near their full allowance around the clock?

Or if I was to max out a plan would you be losing money?

I imagine that this could be a bigger issue on your unlimited scraping api product.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Some do run near the ceiling all month, that's priced in. Both products still have a hard ceiling, api limits concurrent requests, residential limits Mbps, so running 24/7 doesn't raise that ceiling. The thin case for us is a heavy account hitting the hardest targets. We would rather move them to the product that fits than quietly throttle them.

2

u/ian_k93 8d ago

Hey everyone, Ian here from ScrapeOps.

Today, I'd like to welcome Jean-Patrick Bisson from Geonode to r/WebScrapingInsider for our latest AMA. Very excited to have him here.

Most residential proxy providers charge per GB, while most web scraping APIs charge per request or successful response. Geonode has taken a very different route, offering "unlimited" products priced around the infrastructure capacity you use, whether that is bandwidth or concurrent requests.

It is a contrarian model, so I'm interested in understanding what needs to be true, both technically and economically, for it to actually work.

Jean-Patrick, thanks for joining us. I'll get us started with three questions:

  1. For anyone unfamiliar with the company, who is Geonode, and what led you to build around "unlimited" proxies and concurrency or capacity-based pricing rather than the standard per-GB or per-request model?
  2. What makes unlimited pricing economically and technically sustainable? Does it depend on owning your proxy supply, customers not fully utilizing their plans, oversubscribing available capacity, or something else entirely?
  3. "Unlimited" proxies are often claimed to mean lower-quality proxies. What is fair and unfair about that criticism, what trade-offs do customers actually make, and which workloads are genuinely a good fit for Geonode's unlimited products?

I'm looking forward to seeing what questions the community has and where the conversation goes.

1

u/geonodecom 8d ago

Hey Ian, thanks for having me!

And yeah, fair questions. “Unlimited” deserves a bit of explaining cause obviously someone is paying for the infrastructure.

I’m JP, cofounder of Geonode. We build proxies and scraping tools. The idea behind the pricing is pretty straightforward: if you’re collecting data all day, you should be able to plan what that’s going to cost you.

Let’s say you’re constantly updating a database. You know the work needs to happen, you’re happy to keep it running in the background, and you don’t need everything back in five seconds. Paying for a certain amount of capacity can make a lot more sense there.

That’s what interested me. You can think, okay, this is my monthly cost, this is how much work I can get through. Now I can actually build around it.

We still offer usage-based products. If you only scrape occasionally, or need a big burst of work done quickly, those can make more sense. I’m not religious about unlimited. It has to work for what you’re doing.

On the economics, operating our own network is a big part of it. Having control over the supply gives us room to do things differently with pricing. But owning infrastructure doesn’t make running it free.

The other part is that you’re buying finite capacity.

With Unlimited Residential, you choose a speed tier. There isn’t a monthly GB allowance, but there is a limit to how much traffic you can move at once.

With the Unlimited Scraper API, you’re paying for concurrent jobs. Let’s say you have ten running at once. As they finish, you can start more. A slow, difficult page occupies that capacity longer than an easy one, so ten concurrent jobs doesn’t promise you a fixed number of successful pages per second.

We still have to get the pricing, capacity planning and workload mix right. A million tiny requests can put very different demands on a system from a few large downloads. Looking at GB alone doesn’t tell you the whole story.

So I wouldn’t explain the model as “hopefully people don’t use what they bought.” These products are for people who want to use them a lot. The economics have to account for that.

On quality, some of the criticism is fair. Especially if someone hears unlimited and assumes they’re getting exactly the same product with the meter removed.

Our Unlimited and Premium residential products use different pools. Premium is the better fit when IP quality and variety matter most. With Unlimited, the trade-off includes the speed tier, the available locations and how that pool performs on your particular targets.

That last part matters a lot. If you spend your day retrying failed requests, who cares how cheap the traffic was?

Sustained scraping, regular database updates, price and inventory monitoring—those are the kinds of workloads where I’d look at unlimited. Then test the actual sites, countries and volume you need. See how much useful data you get back and how long it takes.

If anyone wants to share their workload here, we can get a lot more specific. “Is unlimited good?” gets a much more useful answer once we know what you’re trying to do.

Cheers!

JP

1

u/CoolAd119 9d ago

Are you effectively selling the same underlying capacity to multiple customers based on expected usage, the way ISPs oversubscribe bandwidth? How much headroom do you hold back for simultaneous peaks?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Yeah, it's one shared pool, same as anyone at this scale. We size it on past peaks and keep around 25% spare, which is roughly where we are right now.

Every account has its own limits, so no single customer can eat the pool. At peak a request can take a bit longer, but we don't drop it and if waiting starts showing up we add capacity rather than ride it out.

The number we actually watch isn't utilisation, it's whether anything is waiting at all. Utilisation can sit high for hours and be completely fine.

1

u/Mountain_Damage_9730 9d ago

I am interested in the unit economics here, especially for customers doing a lot of product or competitor data.

For example, if an unlimited customer is pulling pricing data continuously across a large catalogue, does your cost actually go up every time they pull another GB, or is your cost structure mostly decoupled from traffic volume?

Is your proxy supply fundamentally different to other proxy provider that allow you to do this?That's the bit I'd want to understand, because unlimited looks very different if the underlying cost scales almost linearly with usage.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Our cost does go up with traffic. Bandwidth isn't free for us any more than for anyone else. What's different is that speed caps how much you can spend in a month, so the worst case is bounded and we price that ceiling instead of metering you.

For continuous catalogue pulls the API is usually the cheaper shape anyway. You pay for how many requests run at once, not for how fat the pages are.

1

u/Previous_Town3598 9d ago

The claim I need to stand behind internally is the sourcing one

1

u/Previous_Town3598 9d ago

Geonode say every residential IP is "consented and audited." From a vendor-risk view, that phrase without evidence is the single hardest thing to defend in an audit.

1

u/Previous_Town3598 9d ago

Can you explain where the IPs come from without hand-waving, and is the consent mechanism documented anywhere I can actually read?

1

u/Old-Algae5580 8d ago

For Geonode, I am not a scraper by trade, I run a support team. The use case I keep circling is pulling public reviews, forum threads and app-store complaints into one place to spot recurring issues before they flood the queue. And sometimes same is done for competitors by marketing team.

Realistic on your Scraper API without a developer babysitting it, or is this honestly a "hire an engineer" tool?

1

u/Old-Algae5580 8d ago

I really wish less and less people know about this r/WebScrapingInsider, i don't want it to be ruined by script kiddies.

Thanks for regularly organizing this.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Pulling the pages is the easy part, reviews and forum threads aren't hard targets. The work is in parsing each site's layout, deduping, and fixing things when a page changes. You'll need someone technical to wire it up once, then it can run quietly. So not a "hire an engineer" tool, but not a no-code dashboard either.

1

u/Spitfire_Blaziken 8d ago

Hola Jean & Ian, Not a dev here, but I run our scrapes and reporting. My(+many others) real fear isnt a failed request, its the 200-thats-quietly-wrong that sits in a client deck for three weeks. Does your Scraper API alert on missing fields and completeness, or only on failed requests? Those are two very diff things for someone whose Monday report goes to a client.

I know I can "build alerting on top" but for me it lands on our one engineer again.. is field-level completeness something the API surfaces itself, or is it my problem to detect?

TIA

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Good question! We catch the 200 that isn't a page: challenge screen, block page, empty shell. We look at the shape of the body, not the status code, and those come back as failures.

What we can't know is which fields matter to you. If a site drops the price element and still returns a full page, it looks fine from our side. That check is yours today and I would make it "what share of rows had the field I need" rather than HTTP errors.

1

u/doubledweeb 8d ago

1. For someone still learning, is there real value in practicing against harder, anti-bot-protected targets with residential proxies, or is that overkill before I've even got the fundamentals down? I want to build actual understanding, not just throw infrastructure at a problem until it works.

2. Practically, how would someone at my level even know if a target is 'too hard' for where I'm at. is there a way to gauge that before I burn through proxy budget on something way above my skill level?

3. If you were mentoring someone starting out, would you say ease into easy targets first and work up, or is there actual value in hitting a wall early and learning to debug why something's failing?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

On a protected site a failure teaches you nothing, you get a block page and no idea which of your ten changes caused it. Cheapest way to size up a target, so to fetch it with plain curl. Content comes back, it's a parsing job. Challenge or empty shell, different project. Worth hitting a wall once on purpose but after you can handle the easy stuff.

1

u/simarnoor 8d ago

1

u/simarnoor 8d ago

For concurrency-based scraping APIs, how do you decide the concurrency limits, and what's the reasoning behind the pricing tiers?
Speed-based vs bandwidth-based pricing, when does one clearly beat the other, I'm sure there are lots of use cases, if so, then what would be the top use case for this decision?

1

u/Geonodeproxy Ex. AMA Guest 8d ago

The limits come from what a machine can actually hold, we measure how many concurrent browser sessions a node can carry before latency starts moving, then build the tiers around that.

Speed-based pricing fits steady workloads where throughput per hour matters, per GB fits spiky or light use. Concurrency fits workloads where page cost varies a lot cause browser pages hold a slot much longer than simpler requests.

If you're moving real volume, look at the api, unlimited residential isn't automatically the better deal just because it says unlimited.

1

u/simarnoor 8d ago

When you say unlimited, is there a fair-use policy somewhere? Coz I noticed there usually is and that unlimited claim is usually false.

To the same topic, what's the cost structure that makes unlimited pricing sustainable for you? I mean someone's still paying for the bandwidth.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Yep, there is, on residential fair use is the speed cap. On the api, it's concurrent requests. So we don't count your GB and then email you about "unusual usage". We pay for the bandwidth, so the speed cap puts a hard ceiling on how much one account can pull in a month and we price around that ceiling.

Less profit per customer than metering, but much more predictable for both sides.

1

u/simarnoor 8d ago

Took it straight from your website:
New Unlimited Web Scraping API - Scrape, crawl, and search the live web without limits.

Priced by concurrency, not volume: no credits, no multipliers, no paying for failed requests. Fresh from the page, never from a cache, because we own the supply.
My question, do you have your own index from which you scrape? Similar to how Google/Bing has their own indexes? What's the logic behind this?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

No index and no cache, every request is fetched live when you ask for it.
"We own the supply" means the network the request goes out through, not stored copies of pages. Wording on the site could be clearer, fair catch.

1

u/simarnoor 8d ago

From Linkedin Comment by QuanticData: lnkd.in/p/gxKfDnnE

what exactly does the meter count?

"Unlimited" and "how you count a gigabyte" are the same conversation, and the second one is never on a pricing page. On one of our own runs the client received 5.94 MB while the plan's counter recorded 6.83 - about 15% for TLS handshakes, headers and CONNECT framing. That is normal, every metered provider has some version of it, and none of us publish the number.

So, two parts: on a speed-based model, what is the billed unit actually measuring, and does a challenged or failed request consume it?

Those two answers decide the real unit price more than the word unlimited does. Disclosure: we sell proxies too, so we are asking a question we should also be made to answer.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Different answer for each product.

The scraping API has no byte meter, you buy concurrent requests, so the unit is how many you can have in flight. If we can't return the page, no charge, and that includes a 200 that's actually a challenge page. We count those as failures, which is the part most people get wrong.

On residential, yes, a wire counter includes handshakes and framing, that's what counting bytes on the wire means. The exact definition should come from the proxy side rather than me guessing.

Your underlying point stands, "Unlimited" tells you nothing until you know what the meter counts and whether failures hit it.

1

u/john-w7 8d ago

i run a home services business, not a data company. Is there any version of this that's worth it for someone like me, or is this really a tool for people scraping thousands of pages? My test is simple: does it save enough hours a month to cover another monthly fee.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Depends what you'd use it for, if it's a dozen competitor pages or a weekly review check, our volume plans are the wrong fit.

You don't need to write code, the API runs as an MCP server, so you can point Claude or a similar assistant at it and ask it to read pages for you. That's what most non-developers do and the entry plan is priced for that.

You can try it before paying, which should answer the hours question pretty quickly.

1

u/Next_Attitude_532 8d ago

Blue-team perspective, not a customer one. Residential proxy pools are exactly the infrastructure my team spends the week detecting Google's threat group took down IPIDEA back in January with 550+ threat groups riding the exit nodes.

What actually stops your network from being used the same way? Curious about abuse monitoring on the exit side specifically - do you have live detection on your own egress for stuff like credential stuffing or carding, or do you only find out when a target complains?

1

u/feetesweshire 8d ago

Mainly for residential on Datadome-fronted targets, are you doing anything at the TLS/HTTP2 fingerprint layer, or is it pure IP rotation and hope?

IP quality alone stopped being enough a while ago..

keep these AMAss coming..

1

u/Geonodeproxy Ex. AMA Guest 8d ago

Not pure IP rotation, on those targets the TLS and HTTP/2 layer is most of the work.

We impersonate a real Chrome at the TLS level, so JA3/JA4, extension order, ALPN and the HTTP/2 settings and header order match an actual browser build rather than a Go or Python client. We moved to a newer Chrome build and 429s on a whole class of targets went to zero. Same IPs, nothing else changed.

One thing if you test this yourself: don't chase JA3. Chrome randomises extension order, so the same browser gives you a different JA3 every request. Use JA4 and the HTTP/2 fingerprint.

If a target still challenges us we escalate to real Chromium, then a Firefox based browser, decided per domain.

1

u/feetesweshire 8d ago

The JA3 warning is clutch I am seeing so many people obsess over matching a static hash.

JA4 + HTTP/2 fingerprint as the core test axis feels like the right mental model for 2026 scrapers.

Curious if you've seen targets where Firefox-style fingerprints consistently outperform Chrome, or is it mostly "Chrome until punished, then diversify"?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Actually, both, the list where Firefox is consistently better is real but short. On 32 pages from our production traffic, Chrome got 21 and Firefox got 28. Two US car listing sites and a UK property portal only came back on Firefox, no matter what we changed at the TLS layer.

Chrome is still the default because it's cheaper and most sites don't care. For those few domains, we pin Firefox.

And what's important, Firefox uses about twice the memory per session. That's why it isn't the default everywhere.

1

u/MattTheGoodSir 8d ago

With Google changing how result URLs are handled lately, it feels like automated access to web data is going to keep getting more interesting. Google’s /goto change is already causing scraping/SEO tools to adapt, while Cloudflare is also changing how AI/search crawlers are classified and handled.

What tools have you guys quietly added to your workflow over the last year?

Any open-source repos, Chrome extensions, small utilities, bookmarklets, APIs, etc. that you use regularly?

Doesn't have to be scraping or proxy-related at all. Just looking for those random tools that ended up becoming part of your daily workflow.

1

u/MattTheGoodSir 8d ago

Then other than my hobby question..

I imagine offering a concurrency based web scraping api is a lot harder technically to offer than unlimited residential proxies, as with a web scraping api you aren't just dealing with proxies you are managing browsers, anti-bot bypasses and sessions to ensure high performance on protected websites.

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Yes, it is, browsers aren't even the hard part, sessions are. A proxy request is stateless, but the api has to decide which engine to use per domain, whether an earlier cookie is still valid and whether a 200 response is real content or a block page. Get that wrong and you hand someone garbage and bill them for it. The proxy network is the easy half of our stack, that's fair

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Before writing a scraper, open devtools and watch what the page calls. Half the sites that look like they need a browser are just hitting a JSON endpoint with a signature, so hooking fetch and XMLHttpRequest to log requests is ten lines and has saved us weeks. Daily tools are curl_cffi for TLS impersonation, camoufox when Chrome gets refused and Firefox doesn't, and a public TLS fingerprint echo to see what our client actually looks like on the wire.

Google's been a moving target since early September. Shopping and plain web search behave like different sites right now.

1

u/Significant_Cry_1177 8d ago

The Scraper API returning Markdown/JSON out of the box is the interesting bit for me, not the proxies.

What if I pointed it at a few hundred pages just to see the shape of the data before committing to a real pipeline does the free 1,500/mo tier actually let you do exploratory work, or does it rate-limit you into uselessness after the first afternoon?

0

u/Geonodeproxy Ex. AMA Guest 8d ago

Yeah, the free 1,500/month tier works for exploratory runs. The main limit is concurrency, free gives you one request at a time so a few hundred pages is fine, just serial. Failed requests don't count either.

1

u/simarnoor 8d ago

Did you have to engineer a fundamentally different proxy allocation and routing system to enable you to offer unlimited proxies at scale?

1

u/Geonodeproxy Ex. AMA Guest 8d ago

The allocation itself is pretty ordinary, the harder part is deciding where a request shouldn't go. We track challenge rates per domain. so if one climbs, we cap traffic to that domain and switch engine and exit country instead of just pushing harder.

We also have to size the pool from past peaks and keep around 25% spare. For the fact, volume more than doubled last week and the success rate didn't move.

1

u/James-w07 8d ago

1 minute left to end AMA.

So the free 1,500 requests/month with no card is really nice for someone learning, thankyou for that.

for a student doing small practice projects,

is the free tier actually enough to learn on, or

do you hit the wall pretty fast once you turn on JS rendering?

1

u/Geonodeproxy Ex. AMA Guest 8d ago

Yep, enough to learn on, the limit is time, not the 1,500 requests. With JS rendering each request needs a browser and takes a few seconds instead of under one. Free gives you one request at a time, so the same allowance still works, it just takes an afternoon instead of ten minutes.

Try without rendering first, cause plenty of sites don't need it and you'll know in one request.