A Scrapy Proxy Middleware That Rotates IPs per Request
The quick-and-dirty way to add proxies to Scrapy is setting meta["proxy"] per request inside start_requests. It works for small jobs, then stops working the moment you enable concurrency: you cannot tell whether the proxy was ever applied, and retries keep hammering the same dead IP. The right place for this is a downloader middleware — it can attach a proxy before dispatch and swap it after a failure.
The two hooks that matter
process_request(request, spider)— runs before the request is sent. Setrequest.meta["proxy"]here.process_exception(request, exception, spider)— runs after the request raised. Pick a new IP, assign it, and return the request; Scrapy reschedules it.
Most people only write the first one, which is exactly why "Scrapy retried three times and still used the dead proxy" happens.
Return values are load-bearing: returning
Nonefromprocess_requestmeans "continue"; returning aRequestreplaces the current one; returning aResponseshort-circuits. Getting this wrong makes requests silently disappear.
Ordering
Lower numbers run first. Placing the proxy middleware after the retry middleware is usually right — let Scrapy exhaust its own retries first, then let us rotate:
DOWNLOADER_MIDDLEWARES = {
"proxy_middleware.XydailiProxyMiddleware": 543,
}
The middleware
Note the trust_env = False line: the extraction request must go direct, otherwise a container-level HTTP_PROXY variable will push it through a proxy too.
import time
import requests
API = "http://api.xydaili.net:2022/tools/ip.ashx"
class XydailiProxyMiddleware(object):
def __init__(self, order, api=API, qty=10, area="", isp="",
pool_size=5, max_age=170):
self.order = order
self.api = api
self.qty = qty
self.area = area
self.isp = isp
self.pool_size = pool_size
self.max_age = max_age
self._items = []
@classmethod
def from_crawler(cls, crawler):
s = crawler.settings
order = s.get("XYDAILI_ORDER")
if not order:
raise ValueError("set XYDAILI_ORDER in settings.py")
return cls(
order=order,
qty=s.getint("XYDAILI_QTY", 10),
area=s.get("XYDAILI_AREA", ""),
isp=s.get("XYDAILI_ISP", ""),
)
def _refill(self):
params = {
"action": "GetIP",
"OrderNumber": self.order,
"protocol": 1,
"qty": self.qty,
"split": "json",
}
if self.area:
params["Area"] = self.area
if self.isp:
params["Isp"] = self.isp
session = requests.Session()
session.trust_env = False # extraction must go direct
try:
data = session.get(self.api, params=params, timeout=10).json()
finally:
session.close()
if data.get("status") != 200 or not data.get("data"):
return
now = time.time()
for item in data["data"]:
self._items.append(
("http://%s:%s" % (item["ip"], item["port"]), now + self.max_age)
)
def _take(self):
now = time.time()
self._items = [it for it in self._items if it[1] > now]
if len(self._items) < self.pool_size:
self._refill()
return self._items.pop(0)[0] if self._items else None
def process_request(self, request, spider):
proxy = self._take()
if proxy:
request.meta["proxy"] = proxy
return None
def process_exception(self, request, exception, spider):
"""Swap the proxy and requeue the request."""
proxy = self._take()
if not proxy:
return None
request.meta["proxy"] = proxy
return request
And the settings it expects:
XYDAILI_ORDER = "YOUR-ORDER-ID"
XYDAILI_QTY = 10 # extract in batches, never one at a time
XYDAILI_AREA = "" # optional city filter
XYDAILI_ISP = "" # optional carrier filter
CONCURRENT_REQUESTS = 16
RETRY_TIMES = 2
DOWNLOAD_TIMEOUT = 15
DOWNLOADER_MIDDLEWARES = {
"proxy_middleware.XydailiProxyMiddleware": 543,
}
Three traps
1. Retries do not change the proxy
Scrapy's built-in retry middleware reschedules the request but leaves meta["proxy"] untouched. RETRY_TIMES alone will happily retry the same dead exit three times. Rotation has to come from process_exception. Verify by logging the proxy next to every retry.
2. Adding proxies made failures go up
Usually a concurrency/pool mismatch. With CONCURRENT_REQUESTS = 32 and a pool that only holds 10 IPs, 32 requests fight over 10 exits and each one gets hammered — which looks exactly like abusive traffic to the target. My rule of thumb: pool size around one third to one half of concurrency, and 2–5 concurrent requests per IP.
3. An empty pool silently disables the proxy
As written above, when _take() returns nothing the request goes out directly — and if the target blocks your own IP, the whole batch fails with no obvious cause in the logs. In production, fail loudly instead:
from scrapy.exceptions import IgnoreRequest
def process_request(self, request, spider):
proxy = self._take()
if not proxy:
raise IgnoreRequest("proxy pool empty, skipping %s" % request.url)
request.meta["proxy"] = proxy
return None
Prove the rotation works
Do this once, before trusting anything:
import scrapy
class IpCheckSpider(scrapy.Spider):
name = "ipcheck"
start_urls = ["http://httpbin.org/ip"]
def parse(self, response):
self.logger.info("egress: %s", response.text)
If the echoed address differs from your own public IP, the middleware is live. Two minutes here saves hours of arguing about whether the proxy applied or the site is blocking you.
Summary: attach a proxy in process_request, rotate in process_exception. Then tune concurrency, pool size and per-IP pressure together — changing one alone rarely helps.
The file lives in the scrapy/ directory of github-xydaili-examples, with the settings documented at the top. The IPs I use come from 星月代理, which hands out short-lived rotating IPs in batches — a good fit for the pool above.