A content scan crawls public pages on your connected domain and imports them as candidate knowledge. Publishing in Knowledge is what makes pages available to the agent in chat.
Separate catalog scans (Products, Blog, Help) look for store or help-center URL patterns. Run the main content scan first unless Documentation tells you otherwise for your stack.
Run a scan
- Open the website in your workspace
- Go to Scan or start a scan from Overview
- Wait until status shows completed or indexed — partial success still imports reachable pages
- Open Knowledge, review imported pages, and publish what should drive answers
- Schedule another scan after large content or URL structure changes
After the scan
- Check Scan for failed URLs and fix reachability or robots blocks
- Leave out drafts, thin stubs and pages you do not want quoted
- Add FAQs or notes in Knowledge for facts your site never states clearly
- Open Preview agent and test questions before you treat answers as live-ready
Why pages skip import
- Login walls and member-only areas are not public — the crawler cannot read them
- robots.txt disallow rules and noindex meta tags exclude URLs on purpose
- Broken redirects or 404s show as failures in Scan — fix the live URL first