A Scrapy Proxy Middleware That Rotates IPs per Request

The quick-and-dirty way to add proxies to Scrapy is setting meta["proxy"] per request inside start_requests. It works for small jobs, then stops working the moment you enable concurrency: you cannot tell whether the proxy was ever applied, and retries keep hammering the same dead IP. The right place for this is a downloader middleware — it can attach a proxy before dispatch and swap it after a failure.

The two hooks that matter

  • process_request(request, spider) — runs before the request is sent. Set request.meta["proxy"] here.
  • process_exception(request, exception, spider) — runs after the request raised. Pick a new IP, assign it, and return the request; Scrapy reschedules it.

Most people only write the first one, which is exactly why "Scrapy retried three times and still used the dead proxy" happens.

Return values are load-bearing: returning None from process_request means "continue"; returning a Request replaces the current one; returning a Response short-circuits. Getting this wrong makes requests silently disappear.

Ordering

Lower numbers run first. Placing the proxy middleware after the retry middleware is usually right — let Scrapy exhaust its own retries first, then let us rotate:

DOWNLOADER_MIDDLEWARES = {
    "proxy_middleware.XydailiProxyMiddleware": 543,
}

The middleware

Note the trust_env = False line: the extraction request must go direct, otherwise a container-level HTTP_PROXY variable will push it through a proxy too.

import time

import requests

API = "http://api.xydaili.net:2022/tools/ip.ashx"


class XydailiProxyMiddleware(object):
    def __init__(self, order, api=API, qty=10, area="", isp="",
                 pool_size=5, max_age=170):
        self.order = order
        self.api = api
        self.qty = qty
        self.area = area
        self.isp = isp
        self.pool_size = pool_size
        self.max_age = max_age
        self._items = []

    @classmethod
    def from_crawler(cls, crawler):
        s = crawler.settings
        order = s.get("XYDAILI_ORDER")
        if not order:
            raise ValueError("set XYDAILI_ORDER in settings.py")
        return cls(
            order=order,
            qty=s.getint("XYDAILI_QTY", 10),
            area=s.get("XYDAILI_AREA", ""),
            isp=s.get("XYDAILI_ISP", ""),
        )

    def _refill(self):
        params = {
            "action": "GetIP",
            "OrderNumber": self.order,
            "protocol": 1,
            "qty": self.qty,
            "split": "json",
        }
        if self.area:
            params["Area"] = self.area
        if self.isp:
            params["Isp"] = self.isp

        session = requests.Session()
        session.trust_env = False        # extraction must go direct
        try:
            data = session.get(self.api, params=params, timeout=10).json()
        finally:
            session.close()

        if data.get("status") != 200 or not data.get("data"):
            return
        now = time.time()
        for item in data["data"]:
            self._items.append(
                ("http://%s:%s" % (item["ip"], item["port"]), now + self.max_age)
            )

    def _take(self):
        now = time.time()
        self._items = [it for it in self._items if it[1] > now]
        if len(self._items) < self.pool_size:
            self._refill()
        return self._items.pop(0)[0] if self._items else None

    def process_request(self, request, spider):
        proxy = self._take()
        if proxy:
            request.meta["proxy"] = proxy
        return None

    def process_exception(self, request, exception, spider):
        """Swap the proxy and requeue the request."""
        proxy = self._take()
        if not proxy:
            return None
        request.meta["proxy"] = proxy
        return request

And the settings it expects:

XYDAILI_ORDER = "YOUR-ORDER-ID"
XYDAILI_QTY = 10          # extract in batches, never one at a time
XYDAILI_AREA = ""         # optional city filter
XYDAILI_ISP = ""          # optional carrier filter

CONCURRENT_REQUESTS = 16
RETRY_TIMES = 2
DOWNLOAD_TIMEOUT = 15

DOWNLOADER_MIDDLEWARES = {
    "proxy_middleware.XydailiProxyMiddleware": 543,
}

Three traps

1. Retries do not change the proxy

Scrapy's built-in retry middleware reschedules the request but leaves meta["proxy"] untouched. RETRY_TIMES alone will happily retry the same dead exit three times. Rotation has to come from process_exception. Verify by logging the proxy next to every retry.

2. Adding proxies made failures go up

Usually a concurrency/pool mismatch. With CONCURRENT_REQUESTS = 32 and a pool that only holds 10 IPs, 32 requests fight over 10 exits and each one gets hammered — which looks exactly like abusive traffic to the target. My rule of thumb: pool size around one third to one half of concurrency, and 2–5 concurrent requests per IP.

3. An empty pool silently disables the proxy

As written above, when _take() returns nothing the request goes out directly — and if the target blocks your own IP, the whole batch fails with no obvious cause in the logs. In production, fail loudly instead:

from scrapy.exceptions import IgnoreRequest

def process_request(self, request, spider):
    proxy = self._take()
    if not proxy:
        raise IgnoreRequest("proxy pool empty, skipping %s" % request.url)
    request.meta["proxy"] = proxy
    return None

Prove the rotation works

Do this once, before trusting anything:

import scrapy

class IpCheckSpider(scrapy.Spider):
    name = "ipcheck"
    start_urls = ["http://httpbin.org/ip"]

    def parse(self, response):
        self.logger.info("egress: %s", response.text)

If the echoed address differs from your own public IP, the middleware is live. Two minutes here saves hours of arguing about whether the proxy applied or the site is blocking you.

Summary: attach a proxy in process_request, rotate in process_exception. Then tune concurrency, pool size and per-IP pressure together — changing one alone rarely helps.

The file lives in the scrapy/ directory of github-xydaili-examples, with the settings documented at the top. The IPs I use come from 星月代理, which hands out short-lived rotating IPs in batches — a good fit for the pool above.