How this was built,
including the bugs

Most tool sites describe what they do. This is a record of how these particular tools came to work, including four things that were wrong in code that appeared to run perfectly. It is here because a claim about correctness is worth more when the failures are published alongside it.

Everything is parsed by hand

There is no third-party library between you and your file on this site. The JPEG segment walker, the TIFF directory traversal that reads EXIF, the HEIC container reader, the PDF object scanner and the ZIP reader were all written for this. Archive compression uses the browser's own DecompressionStream and CompressionStream rather than a bundled compressor.

That was a deliberate trade. A library would have been faster to write and would cover more edge cases. It would also mean asking people to trust code neither of us has read, on a site whose entire proposition is that you do not have to take anything on trust. One readable page is worth the extra work.

Four bugs that testing caught

1. Photographs came out rotated twice

Images with an EXIF orientation tag were being turned the wrong way when converted to PDF. The cause was that createImageBitmap in Chrome already applies the orientation tag while decoding, and the code then applied it again on top. A photograph taken sideways came out sideways in the other direction.

The fix is not to pick one behaviour and assume it, because browsers differ. The decoded dimensions are now compared against the dimensions written in the JPEG header. If they match the stored values, the tag was ignored and the rotation is applied; if they are swapped, the browser has already done it and nothing further is needed.

2. A missing namespace broke every spreadsheet

Cleaning a Word document worked. Cleaning an Excel workbook produced a file that Excel refused to open, with a complaint about an undefined namespace prefix.

The reason is a difference in how two tools write the same XML. One declares its namespaces on the root element, the other declares them inline on the elements that use them. The cleaning code blanked an element by replacing it wholesale, which silently discarded the namespace declaration attached to it. Word documents survived; spreadsheets did not. The code now preserves each element's own attributes while emptying its contents. This was only found because the tests reopened every cleaned file rather than checking that a file had been produced.

3. The accessibility checker missed the thing it was checking for

The WCAG check for a visible focus indicator looks for a stylesheet that removes the outline without providing a replacement. Given a page containing exactly that mistake, it reported nothing.

The guard clause was the problem. It skipped the finding if the stylesheet contained a focus rule mentioning an outline — and *:focus { outline: none } is a focus rule mentioning an outline. The check was being disarmed by the very code it existed to catch. It now looks for a focus rule that sets an indicator to something other than none, zero or transparent.

4. Crawlers were requesting pages that do not exist

The server logs filled with 404s for addresses ending in ${l.url} and ${url}. Those are JavaScript template placeholders, and they were appearing in the raw HTML because the markup for the report was written as a template string containing src="${url}". At runtime the value fills in correctly. In the page source, before any script runs, the literal text sits there — and crawlers scrape raw HTML for src and href attributes.

Harmless to a visitor, wasteful for a new site with a limited crawl budget. Those attributes are now written as data-src and promoted to real ones after the markup is inserted.

Decisions that look like limitations

Hidden spreadsheet sheets are reported and left alone

A hidden sheet usually contains the working behind the visible numbers, and formulas on the visible sheets frequently reference it. Deleting it would leave a recipient looking at reference errors instead of results, which is worse than the disclosure. It is flagged clearly and left in place, because destroying data is not a decision a tool should make quietly.

Some PDFs cannot be fully cleaned

A PDF saved incrementally keeps its earlier versions in the document body rather than in its metadata. Removing them means genuinely rewriting the file, which is a different and riskier operation. The checker detects this, reports it as its own finding, and tells you to re-export the document instead of implying the problem is solved.

A URL cannot be checked for accessibility

Browsers do not allow one site to read another site's pages, and the workaround — sending the address to a server that fetches it — is exactly what this site does not do. So the accessibility checker takes pasted markup instead. It is less convenient and it keeps the promise intact.

What is tested, and how

Every format is checked against files with known contents planted in them: a photograph carrying coordinates, a camera serial and an owner name; a Word document with tracked changes and two named editors; a spreadsheet with a hidden sheet; a PDF saved several times; an archive containing an environment file, a private key and a database backup. Each is run through the parser, cleaned, and then reopened with independent software to confirm it still works and that the planted data is gone. The cleaned image is compared pixel by pixel against the original to confirm nothing was re-compressed.

The accessibility checker is tested in both directions, which matters more than it sounds: against a deliberately broken page, where it should find everything, and against a well-built page, where it must find nothing. A checker that reports false problems gets ignored, and an ignored checker is worse than no checker.

The tools