Repository navigation
Fix unbounded sitemap discovery on narrow base URLs - #140
Merged
Merged
Conversation
Bound sitemap walks independently of accepted URLs, add a cumulative body budget and narrow-prefix diagnostics, and reuse the bounded walker for coverage fallback. Add regression tests and document request-budget semantics for #120.
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #120.
The sitemap walker only stopped after collecting enough accepted URLs. If a base-path filter rejected every page, it could fetch every shard in a large index. The locale filter also missed Microsoft-style filenames such as
dotnet_en-us_1.xml.Budget Semantics
Limits apply per walk, not per scan. Robots discovery and existing HTTP redirects/retries are separate from the 20 sitemap fetch calls. No per-page probes or deeper traversal are added, and existing sample and URL collection limits remain unchanged.
The byte limit is checked between completed responses: the final body can overshoot it, and failed partial reads do not expose a byte count. This is not a hard streaming memory limit or a wall-clock guarantee. Existing request/body timeouts still apply. Coverage retains budget warnings in its result details; truncated results are not exhaustive site coverage.
Verification
The initial two regressions failed before the fix: the filter retained the French shard, and the all-rejected walk fetched all 26 fixture sitemap documents instead of stopping at 20.
npm run lint,npm run format:check,npm run version:check, andnpm run buildpassed.Local validation used Node 25.2.1; CI covers Node 22 and 24. The original Microsoft crawl was rerun successfully; see the live results below.
Live Verification (2026-09-27)
Rebuilt commit
2a09a73and ran the existing field harness against the exact saved baseline URL,https://learn.microsoft.com/en-us/docs/, using the same deterministic 20-page sampling, 200 ms request delay, and 12-minute safety cutoff.All nine sitemap responses were HTTP 200 with no recorded errors. The eight fetched shards were
previous-versions_en-us_1.xmlthroughprevious-versions_en-us_8.xml; no other locale shards were fetched. Discovery emitted:Sitemap body byte limit (52428800 bytes) reached; discovery results are partial.Sitemap path prefix "/en-us/docs" matched 0 of 329018 examined same-site URLs (less than 1%); consider a broader base URL.The live run reached the byte budget before the 20-fetch ceiling. It confirms the locale filtering, budget stop, narrow-prefix diagnostic, and completion of the scan with the existing single-base-page fallback. The independent fetch ceiling remains covered by regression tests. Coverage was skipped because no llms.txt was found, so the coverage fallback was not exercised live.
This is a reproduction check, not a controlled performance benchmark: site content and network conditions can change. Exact bytes beyond the threshold are not retained by the field ledger. The original ledger remains untouched at
bot-results/microsoft.json; the fixed ledger is saved locally atbot-results/microsoft-issue120-fixed-20260927.json(both gitignored).