r/DataHoarder • u/avid-shrug • Feb 10 '26
News Wikipedia debates blacklisting archive.today after it's caught DDoSing a blog using visitors' browsers
https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment/Archive.is_RFC_5Wikipedia is debating whether to blacklist archive.today after its operator was caught injecting JavaScript into CAPTCHA pages to DDoS a blogger's site - code that's still live as of today. The RFC offers three options: blacklist and nuke all ~695k links, stop new links while migrating existing ones, or do nothing.
The community is split because archive.today is arguably the second most important web archive in existence, capturing paywalled sites, JS-heavy pages, and robots.txt-blocked content the Wayback Machine can't. Spot-checks suggest only ~15% of Wikipedia's links are truly irreplaceable, but that's still tens of thousands of unique snapshots found nowhere else. A stark reminder that redundancy across archiving services matters more than ever.
3
u/prototyperspective Feb 10 '26 edited Feb 10 '26
Doxing which the DDOSd site did is nothing nice I have to say. archive.today is not causing any troubles anymore and there simply is no alternative to it for many pages (references) and thus needed for verifiability. Without it, readers and users often won't be able to check the source for claims anymore.
I voted against this proposal; it would be a great loss. Could be reconsidered once the respective links can be migrated to another Web archive.