Found via CSS background-image attributes in WBM snapshots (missed
by previous scripts that only checked src/data-src). Also adds
recover_content.py refinements. Remaining ~93 images (verduras 04-27,
arroz negro steps, arrozabanda/fideua paso series) confirmed absent
from all public archives — never crawled by WBM/CommonCrawl/arquivo.pt.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Downloads images from multiple Wayback Machine snapshot dates using
page-crawl extraction (im_ modifier). Adds recover_assets3.py and
recover_content.py for targeted multi-date recovery.
Remaining ~107 content images (2021 step-by-step recipe photos for
verduras, arroz negro, arrozabanda, fideua) are confirmed absent from
all web archives — pages were crawled but images were lazy-loaded.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Downloads 28 previously missing images from Wayback Machine using
im_ modifier and CDX-lookup timestamps. Adds recover_assets.py
(CDX-based) and recover_assets2.py (page-crawl-based) for continued
recovery when WBM rate limit lifts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>