pstoh
Backend / data collection engineer · proxy networks and crawler reliability
I write this site for a simple reason: most material about scraping stops at "install the library". What actually blocks people is the second week — the job that worked yesterday and fails today. These notes are the debugging paths I wish I had found back then.
What I write about
- The proxy layer — how HTTP proxies forward and authenticate requests, pool design, rotation strategy;
- Engineering — Scrapy middleware, rate limiting and backoff, retry policy, failure metrics;
- Post-mortems — the order in which to check things when an error appears, including the cases that look like network problems but are not.
Where the code lives
Every example in the posts is packaged as a runnable project covering Python, Node.js, TypeScript, Java, Kotlin, Scala, C#, VB.NET, Go, PHP, Ruby, Perl, Rust, Swift, Dart, C++, Shell and PowerShell — 18 languages, each showing the same loop: extract an IP, fetch through it, swap and retry on failure.
👉 github.com/pstoh/github-xydaili-examples
About the proxy service
The examples use domestic Chinese HTTP proxy IPs from 星月代理 (Xingyue Proxy) — 300+ cities, short-lived IPs rotated automatically, extraction via API, plus an IP allowlist mode. The measurements in the posts come from its lines. To be clear: this is a technical demonstration, not a purchase recommendation — evaluate any provider against your own workload.
Get in touch
- GitHub: github.com/pstoh
- Bug reports and corrections: open an issue on the repository — easier to track than email
License
Posts are licensed CC BY 4.0; please link back if you republish. Code samples are free to use commercially, no attribution needed.
The short version: proxies are one link in the chain. What decides whether a crawler survives for months is rate limiting, retries, caching and monitoring — those matter far more than how big your IP pool is.