context.route() (request interception) breaks URL-embedded proxy auth in
Chromium — DataImpulse proactive CONNECT auth stops working when interception
is active. Replace route handler with --blink-settings=imagesEnabled=false
Chrome flag to block images without touching the interception layer.
Restore DataImpulse URL-embedded credentials (unchanged from before IPRoyal work).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
cf.bstatic.com (images + CSS CDN) was consuming 1.5 GB per scrape run.
Added route interception to abort image, stylesheet, font, and media
resources, plus pure tracking domains. The DOM extraction only reads
HTML attributes — no visual resources are needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
DataImpulse endpoints are intermittently broken — some return EOF after
CONNECT immediately, others work fine. DNS round-robins between them so
the same request can fail 3× then succeed on the 4th. Timeouts were
returning blocked=False and silently failing with no retry. Now treated
as retryable blocks (with context rotation) so a fresh endpoint is tried.
Max retries 2→4 (5 total attempts) to give enough chances to land a
working DataImpulse endpoint.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Old Stocks block_id has 0 in position 2 (not the persons count) so all
its rates were stored as max_persons=0 and deprioritised in the matrix.
Now reads max_persons from the visible 'Max persons: N' span (matching
the original scrapy spider), with aria-label/title fallback. block_id
segment 2 is unreliable across different hotel layouts.
Also treat max_persons=0 as unknown in the matrix ORDER BY priority.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add UA rotation (8 real Chrome UAs), random viewport, en-GB locale,
Europe/London timezone, stealth browser args to suppress automation
signals, and a per-context init script masking navigator.webdriver,
plugins, and chrome runtime.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The previous approach walked siblings from [id^="room_type_id_"] anchors
and only processed rows with js-rt-block-row class, causing the first rate
row of some room types to be silently skipped when that class was absent.
Now iterates all #available_rooms tbody tr rows with data-block-id (same
strategy as the original scrapy spider), using the room_type_id_ element
presence to identify room names only on the first row, with fallback to the
stored name for subsequent rows of the same room type.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_proxy_enabled() was checking cfg.get('server') but load_config() returns
{host, port, username, ...} — no 'server' key. Server URL is built by
proxy_util.playwright_proxy(). This meant proxy was silently disabled
and _proxy_kwargs() would also KeyError if called.
- _proxy_enabled() now delegates to proxy_util.is_enabled() (checks host+username)
- _proxy_kwargs() now calls proxy_util.playwright_proxy() with credentials
embedded in the URL (avoids 14s DataImpulse 407 round-trip)
- _get_context() generates a sticky session_id per context so each worker
gets a distinct residential IP lane
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Scraper: extract max_persons from block_id (3rd segment) instead of cell text; cell[0] was picking up the price value (e.g. 189/166) causing the max_persons=2 filter to exclude all scraped rows and show "No rates available" in snapshot
- MarketView: add namelength: -1 to Plotly legend to prevent label truncation; increase bottom margin from 50→65 and legend y from -0.25→-0.3
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 'We have X left' scarcity indicator only appears when availability is
low (typically ≤5). The quantity <select> dropdown in the booking form
always shows 0..N where N = actual available qty (capped at 10).
Extract max option value from the <select> in the last table cell per rate
plan row, falling back to regex parsing of the cell text if the select
isn't rendered, then to the scarcity text as a last resort.
Stored in both rooms_left and available_qty on RateData.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Switches booking.com scraping from search-results page (Cloudflare-targeted)
to individual hotel property pages (not CF-protected). Each page load returns
all room types, all rate plan variants (room-only/B&B × refundable/non-ref ×
1-2 adults), and availability counts.
Key changes:
- New PlaywrightHotelPageBackend: proxy reuse until block, rotate on CF/WAF
- booking_scraper.py: _run_hotel_page_scrape(), scrape_hotel_date(),
_scrape_hotels_concurrent() — sharded by hotel so one proxy session covers
all dates for one hotel (looks human)
- schema.sql: ADD COLUMN rate_plan_id, max_persons on booking_com_rates
- competitors API: filter max_persons=2, order by rate_gross ASC as tiebreaker
so DISTINCT ON returns cheapest 2-adult rate from latest batch
Enable via Settings → Scraper Backend → playwright_hotel_page
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When page.goto() times out while Cloudflare's JS challenge redirect is in
progress, page.content() silently blocks until that navigation completes
before raising — adding ~1m46s of hidden delay per blocked-page attempt.
Fix: skip page.content() entirely when page_loaded=False (goto already
timed out), setting content="" so block detection treats it as no signal.
Also break out of the page loop early when page 1 failed to load with 0
hotels — no point spending another 60s on the offset-URL page 2 when the
browser is in a Cloudflare challenge state.
Per-blocked-attempt time: ~4 min → ~1 min (30s goto + 30s selectors).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Timeouts were set for the old proxy 407 round-trip world (14.5s latency
per challenge). Now that credentials are embedded in the proxy URL, pages
load in ~7s — the 90s goto timeout was forcing 5+ minutes of wasted wait
on Cloudflare-blocked dates before rotating sessions.
homepage warmup goto: 90s → 20s
page 1 search goto: 90s → 30s (4× normal load time headroom)
property-card waits: 30s → 15s (cards appear in <3s on clean pages)
Default concurrency 3 → 6: each worker uses its own residential proxy IP
so parallelism is safe. ~0.4 GB per Chromium; 6 workers fits in 4 GB LXC.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
page.content() throws when page is still navigating after a goto timeout;
catch it and continue with empty content. Also broaden the soft-block check
to trigger IP rotation whenever no hotels are found (not just on explicit
blocks), so proxy rotation fires even when the error path is hit.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When all 3 proxy rotation attempts are soft-blocked (bad IP pool), fall
back to the server's own IP for a final attempt. Keeps proxy as primary
for IP diversity / rate-limit protection while ensuring a clean fallback
when the assigned residential pool is Cloudflare-challenged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A direct hit to searchresults.en-gb.html with no Referer causes Cloudflare
to hold the response body open (never sending HTML), so domcontentloaded
never fires. Setting referer=HOMEPAGE makes the request look like the user
searched from the homepage, which resolves the stall.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Headless mode gets blocked by Cloudflare on the search endpoint. Headed
mode with Xvfb works on dev (Unraid). The missing piece on Proxmox LXC was
insufficient /dev/shm (Docker default 64 MB); setting shm_size: 256m gives
Xvfb enough shared memory to start.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
DataImpulse residential proxy takes ~29s to route through Booking.com's
Cloudflare edge — the previous 30s timeout was too short and the page
never had a chance to deliver its body. Also raised wait_for_selector
from 15s to 30s so property cards have time to render post-load.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Xvfb exits immediately on the production LXC (no framebuffer device
support), so headed mode was never working. Switch to headless=True
and remove Xvfb from the Dockerfile entirely.
Stealth coverage is unchanged: playwright-stealth patches webdriver,
plugins and chrome.runtime; the residential proxy + sticky session
provide the GB residential fingerprint; homepage warmup carries real
session cookies into the search. --enable-unsafe-swiftshader and
--disable-blink-features=AutomationControlled are kept.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Proxy config, DataImpulse sticky-session username syntax, and Playwright/
httpx proxy builders now live in one place (services/proxy.py) instead of
being duplicated across the Booking.com backend and the /config/proxy test
endpoint. Both scrapers consume it.
- services/proxy.py: load_config/normalize (DB-authoritative, env fallback),
new_session_id, username, playwright_proxy, httpx_proxy_url
- PlaywrightLocalBackend delegates proxy building to the module
- get_scraper_backend factory uses proxy.load_config (one resolution path)
- test_proxy_config endpoint uses the shared URL builder; httpx proxies= ->
proxy= (forward-compatible, 0.28-safe)
- Direct booking-engine scraper (httpx) can now route through the same proxy,
gated by the direct_scraper_use_proxy flag (default off, plumbing ready)
Dead code removed: set_scraper_paused (never called — rotate-on-block
replaced pause-on-block), get_competitor_matrix / get_hotels_list /
update_hotel_tier (endpoints have their own SQL), unused PROXY_KEYS tuple.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The location_search_url column existed but was dead — the scraper always
rebuilt the URL from the free-text location name. Now the Location
Configuration form takes a "Booking.com search URL" field: paste the
address-bar URL from a real search and the scraper lifts ss/dest_id/
dest_type from it (the most reliable destination pin). Falls back to
dest_id, then plain name. Server also extracts dest_id from the URL for
the column and derives a display name from ss when none is typed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Free-text ss= destination resolution is non-deterministic: tonight's
30-day run resolved "Stow on the Wold" to St. Wolfgang, Austria for 8 of
30 dates, saving Salzkammergut hotel rates into the matrix. dest_id +
dest_type=city in the search URL pins the destination.
- booking_scrape_config gains a dest_id column (idempotent ALTER)
- scrape_location_search/_build_search_url thread dest_id through
- /config/location accepts dest_id
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Route through DataImpulse residential proxy when BOOKING_PROXY_* env
vars are set (empty = direct connection, unchanged behaviour)
- Sticky one IP per session via DataImpulse sessid; rotate_session()
burns the IP for a fresh one
- Rotate-on-block: retry a date on a new IP when the page is challenged
or page 1 renders zero hotels (up to 3 IPs; single attempt sans proxy)
- Persistent, pre-warmed context across dates so cache stays hot
(~5 MB first page, ~0.3 MB per date after) instead of per-date cold loads
- Block images/media/fonts and third-party ad/consent/analytics hosts to
cut bandwidth; never touch the anti-bot challenge (awswaf)
- Faster inter-page pacing now that a burned IP is cheap
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Tier 1+2 anti-detection to stop Booking.com throttling page 2:
- Run Chromium headful under xvfb (Dockerfile) — headless leaks SwiftShader
WebGL, empty plugins, missing chrome.runtime
- playwright-stealth patches navigator.webdriver/plugins/WebGL vendor
- Single coherent Chrome-121 identity: UA + matching sec-ch-ua client hints +
platform (dropped the Firefox/Safari UA strings — a mismatched UA on a
Chromium engine is a stronger tell than no rotation)
- Homepage warm-up so search requests carry real session cookies + cookie
consent dismiss
- Paginate by clicking next (offset= deep-link was the page-2 tell), offset
URL kept as fallback
- Viewport jitter, mouse movement, longer scroll, selector retry on lazy load
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A page-load timeout was logged and skipped, so a scrape could 'succeed'
with only page 1 of results — and the not_listed flagging then marked
every page-2 hotel absent, suppressing their last known rates.
- ScraperResult now tracks pages_requested/pages_ok
- scrape_date skips not_listed flagging when pages failed or when the
scrape saw <60% of the date's 7-day coverage baseline; scraped rates
are still saved, unseen hotels keep last known rate + scrape time
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>