Skip to content

Fix unbounded sitemap discovery on narrow base URLs - #140

Merged
dacharyc merged 2 commits into
mainfrom
fix/120-bounded-sitemap-discovery
Sep 27, 2026
Merged

dacharyc merged 2 commits into
mainfrom
fix/120-bounded-sitemap-discovery

Conversation

@dacharyc

@dacharyc dacharyc commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Summary

Fixes #120.

The sitemap walker only stopped after collecting enough accepted URLs. If a base-path filter rejected every page, it could fetch every shard in a large index. The locale filter also missed Microsoft-style filenames such as dotnet_en-us_1.xml.

  • Recognize underscore/hyphen-delimited filename locales with optional numeric shards, matching URL pathnames independently of query strings. Preserve existing locale validation, two-locale detection, and fallback behavior.
  • Limit each sitemap walk to 20 sitemap fetch attempts across roots and children, including failed and empty responses.
  • Stop subsequent fetches after completed sitemap bodies total 50 MiB of decoded bytes, preserving already-collected URLs.
  • Report partial discovery when a budget prevents another fetch.
  • Warn when a non-root path prefix matches fewer than 1% of examined same-site URLs in sitemap or llms.txt discovery, including the prefix, counts, and broader-base guidance.
  • Reuse the bounded walker for the coverage docs-specific sitemap fallback, which otherwise retained an independent unbounded loop.
  • Record decisions and rejected alternatives in the discovery notes and document CLI behavior.

Budget Semantics

Limits apply per walk, not per scan. Robots discovery and existing HTTP redirects/retries are separate from the 20 sitemap fetch calls. No per-page probes or deeper traversal are added, and existing sample and URL collection limits remain unchanged.

The byte limit is checked between completed responses: the final body can overshoot it, and failed partial reads do not expose a byte count. This is not a hard streaming memory limit or a wall-clock guarantee. Existing request/body timeouts still apply. Coverage retains budget warnings in its result details; truncated results are not exhaustive site coverage.

Verification

The initial two regressions failed before the fix: the filter retained the French shard, and the all-rejected walk fetched all 26 fixture sitemap documents instead of stopping at 20.

  • 233 focused discovery/coverage tests passed, covering locale filename forms, shared budgets across roots/indexes, byte accounting, failed/empty fetches, raw coverage mode, fallback warning propagation, normal termination, and the strict 1% threshold.
  • Full suite: 1,792 tests passed across 68 files.
  • npm run lint, npm run format:check, npm run version:check, and npm run build passed.

Local validation used Node 25.2.1; CI covers Node 22 and 24. The original Microsoft crawl was rerun successfully; see the live results below.

Live Verification (2026-09-27)

Rebuilt commit 2a09a73 and ran the existing field harness against the exact saved baseline URL, https://learn.microsoft.com/en-us/docs/, using the same deterministic 20-page sampling, 200 ms request delay, and 12-minute safety cutoff.

Measurement Saved baseline (2026-09-26) Fixed run (2026-09-27)
Completion Cut off at 12 minutes Finished in 9.561 seconds
Total requests 651 45
Sitemap requests 647 (index + 646 shards) 9 (index + 8 shards)
Check results recorded 5 before interruption All 28, including skips

All nine sitemap responses were HTTP 200 with no recorded errors. The eight fetched shards were previous-versions_en-us_1.xml through previous-versions_en-us_8.xml; no other locale shards were fetched. Discovery emitted:

  • Sitemap body byte limit (52428800 bytes) reached; discovery results are partial.
  • Sitemap path prefix "/en-us/docs" matched 0 of 329018 examined same-site URLs (less than 1%); consider a broader base URL.

The live run reached the byte budget before the 20-fetch ceiling. It confirms the locale filtering, budget stop, narrow-prefix diagnostic, and completion of the scan with the existing single-base-page fallback. The independent fetch ceiling remains covered by regression tests. Coverage was skipped because no llms.txt was found, so the coverage fallback was not exercised live.

This is a reproduction check, not a controlled performance benchmark: site content and network conditions can change. Exact bytes beyond the threshold are not retained by the field ledger. The original ledger remains untouched at bot-results/microsoft.json; the fixed ledger is saved locally at bot-results/microsoft-issue120-fixed-20260927.json (both gitignored).

Bound sitemap walks independently of accepted URLs, add a cumulative body budget and narrow-prefix diagnostics, and reuse the bounded walker for coverage fallback.

Add regression tests and document request-budget semantics for #120.
@dacharyc
dacharyc merged commit 0eee614 into main Sep 27, 2026
3 checks passed
@dacharyc
dacharyc deleted the fix/120-bounded-sitemap-discovery branch September 27, 2026 19:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Page discovery: unbounded sitemap-index walk on narrow base URLs; locale filter misses infix locale names

1 participant