Crawl
Crawl automatically discovers and scrapes multiple pages from a website. Instead of manually specifying each URL, crawl intelligently finds pages and extracts content from each page. Discover pages via sitemaps and links, then get every page as Markdown, HTML or text. Control scope, depth and filters.
How It Works
- Enter Starting URL: Provide the webpage where the crawl should begin.
- Configure Settings: Set up your crawl preferences.
- Automatic Discovery: The robot finds pages using sitemaps and links.
- Content Extraction: Each discovered page is visited and scraped.
What Gets Extracted
For each page discovered, crawl extracts
- Page metadata: Title, language, description, favicon, and all meta tags
- HTML content: Full page HTML
- Text content: Clean body text with word count
- Links: All links found on the page
- Status information: HTTP status code and scrape timestamp
- Screenshots: Visible viewport or full-page screenshot per crawled page
- Summary: A concise plain-text summary of the page generated by an LLM
Basic Configuration
Robot Name
- Custom name for your crawl robot
Starting URL
- The webpage where the crawl begins
- Must be a valid, accessible URL
- All discovered pages will be relative to this starting point
Max Pages to Crawl
- Maximum number of pages to crawl
- Pages are discovered and crawled in order of relevance
- Recommended: Start with 10-20 pages to test your configuration
Advanced Options
Crawl Scope
Choose how broadly the robot should crawl from your starting URL:
Domain Mode
- Crawls only pages on the exact same domain
- Example: Starting at
blog.example.comstays onblog.example.com - Best for: Focused crawls on a specific subdomain
Subdomain Mode
- Crawls the domain and all its subdomains
- Example: Starting at
example.comincludesblog.example.com,shop.example.com, etc. - Best for: Comprehensive site crawls across multiple sections
Path Mode
- Crawls only pages under the same path as the starting URL
- Example: Starting at
example.com/blog/stays within/blog/path - Best for: Crawling specific sections like documentation or blog categories
Max Depth
Maximum Crawl Depth
- Controls how many "levels" deep from the starting URL the crawler should go
- Each click or navigation from one page to another counts as one level
- Higher depth values discover more pages but increase crawl time
URL Filtering
Note: URL Filtering is currently in development and not fully enforced.
Include Paths
- Regex patterns for URLs to include in your crawl
- Only URLs matching these patterns will be crawled
- Leave empty to include all URLs within your scope
- Example:
/blog/[0-9]{4}/.*for dated blog posts
Exclude Paths
- Regex patterns for URLs to exclude from your crawl
- URLs matching these patterns will be skipped
- Example:
.*/admin/.*to skip admin pages - Useful for avoiding login pages, tag pages, or duplicate content
Pattern Tips:
- Patterns use JavaScript regular expressions (regex)
- The
.*matches any characters (wildcard) - Use
\\.to match literal dots in URLs - Patterns are case-sensitive by default
Discovery Options
Use Sitemap
- When enabled, fetches and parses the website's sitemap.xml
- Automatically follows nested sitemaps
- Recommended for sites with well-maintained sitemaps
Follow Links
- When enabled, extracts all links from each visited page
- Crawls pages that match your scope and filters
- Recommended when sitemap is incomplete or unavailable
Best Practice: Enable both options for comprehensive discovery. The robot will combine URLs from both sources and remove duplicates.
Robots.txt Compliance
When enabled, the crawler respects robots.txt directives. Recommended for ethical crawling of third-party websites. By default enabled.
Example Configurations
Simple Blog Crawl
{
mode: 'path',
limit: 50,
useSitemap: true,
followLinks: true
}
Filtered Blog Crawl (Advanced)
{
mode: 'path',
limit: 50,
includePaths: ['/blog/[0-9]{4}/.*'],
excludePaths: ['.*/tag/.*', '.*/author/.*'],
useSitemap: true,
followLinks: true
}
Full Site Crawl
{
mode: 'subdomain',
limit: 100,
excludePaths: ['.*/admin/.*', '.*/login.*'],
useSitemap: true,
followLinks: true
}
Documentation Crawl
{
mode: 'path',
limit: 200,
useSitemap: true,
followLinks: false
}
When to Use Crawl
- You need to scrape multiple pages from a website
- You want to discover pages automatically without listing URLs manually
- You're extracting similar content across many pages (blog posts, product pages, documentation)
- The website has a clear structure or sitemap
When Not to Use Crawl
- You only need data from a single page (use Extract or Scrape instead)
- You need complex interactions like logins or form submissions
- You need to extract structured data in a specific format (use Extract)
- You need to control the exact order pages are visited
For complex workflows with user interactions, use Extract instead.
Using with SDK
Crawl is available through the Maxun SDK for programmatic usage and integration into your applications.
- Node.js
- Python
import { Maxun } from 'maxun-sdk';
const maxun = new Maxun();
const robot = await maxun.crawl('Example docs', 'https://docs.example.com', { limit: 20 });
const result = await robot.run();
for (const page of result.crawlData) {
console.log(page.metadata.url);
console.log(page.markdown?.slice(0, 200));
}
from maxun import Maxun
async with Maxun() as maxun:
robot = await maxun.crawl("Example docs", "https://docs.example.com", limit=20)
result = await robot.run()
for page in result.crawl_data:
print(page["metadata"]["url"])
print(page.get("markdown", "")[:200])
Using with CLI
Crawl is available through the Maxun CLI for quick data gathering from the terminal.
maxun robots crawl https://docs.example.com --limit 20 --include "/docs/*" -n "Docs Crawler"