Skip to main content

Crawl

A crawl robot starts on one page, follows the links it finds and collects the content of every page it visits. Use it to grab a whole documentation site, a blog or a product catalogue.

const robot = await maxun.crawl('Example docs', 'https://docs.example.com', { limit: 20 });

const result = await robot.run();

for (const page of result.crawlData) {
console.log(page.metadata.url);
console.log(page.markdown?.slice(0, 200));
}

Options​

Every option has a sensible default, so a name and a URL are enough to start.

OptionDefaultDescription
mode'domain'Which links to follow: 'domain', 'subdomain' or 'path' (see below)
limit50Maximum number of pages to crawl
maxDepth3How many links away from the start page to go
includePathsnoneOnly crawl URLs matching these patterns, e.g. ['/blog/*']
excludePathsnoneSkip URLs matching these patterns, e.g. ['/admin/*']
useSitemaptrueAlso find pages through the site's sitemap.xml
followLinkstrueFollow links found on each page
respectRobotstrueObey the site's robots.txt
formats['markdown']What to capture from each page: markdown, html, text, links, summary, screenshot-visible, screenshot-fullpage
monitorfalseCompare every run with the previous one. See Monitoring
const robot = await maxun.crawl('Docs guides', 'https://docs.example.com', {
limit: 100,
maxDepth: 4,
includePaths: ['/guides/*'],
excludePaths: ['/guides/archive/*'],
formats: ['markdown', 'links'],
});

Crawl modes​

mode decides how far the crawl may wander from the start URL.

ModeStarting at https://example.com/blog it crawls
domainAny page on example.com
subdomainPages on example.com and its subdomains, such as docs.example.com
pathOnly pages under example.com/blog
const robot = await maxun.crawl('Blog', 'https://example.com/blog', { mode: 'path' });

Reading the results​

result.crawlData is an array with one entry per page:

const result = await robot.run();

for (const page of result.crawlData) {
console.log(page.metadata.url);
console.log(page.metadata.title);
console.log(page.markdown);
}

Each page contains the formats you asked for (markdown, html, text, links, ...) plus metadata with the page's url, title and other meta tags. A page that could not be loaded has an error instead.

tip

Large crawls can take a while. run() waits until the crawl finishes. Pass { timeout: 600000 } (milliseconds) to stop waiting sooner; the crawl keeps going on Maxun and you can read it later with await robot.getLatestRun().

Examples​

A whole documentation site​

const robot = await maxun.crawl('Product docs', 'https://docs.example.com', {
limit: 200,
maxDepth: 5,
});

const result = await robot.run();

const docs = Object.fromEntries(
result.crawlData
.filter((page) => page.markdown)
.map((page) => [page.metadata.url, page.markdown]),
);

Only blog posts​

const robot = await maxun.crawl('Blog posts', 'https://example.com', {
includePaths: ['/blog/*'],
excludePaths: ['/blog/tag/*', '/blog/page/*'],
limit: 50,
});

Keep a knowledge base up to date​

Crawl once a week and see which pages were added, removed or changed:

const robot = await maxun.crawl('Help center', 'https://help.example.com', { limit: 100, monitor: true });
await robot.schedule({ runEvery: 1, runEveryUnit: 'WEEKS', startFrom: 'MONDAY', atTimeStart: '06:00' });

See Monitoring for how to read the changes.

Summarize every page​

const robot = await maxun.crawl('Docs summaries', 'https://docs.example.com', {
formats: ['summary'],
limit: 20,
});

const result = await robot.run();
for (const page of result.crawlData) {
console.log(page.metadata.url, '→', page.summary);
}

summary uses an LLM. On self-hosted Maxun, also pass llmProvider, llmApiKey and optionally llmModel and llmBaseUrl. On Maxun Cloud, leave them out.

Managing crawl robots​

const robots = await maxun.crawl.list(); // all crawl robots
const robot = await maxun.robots.find('Blog posts');
await robot.setListLimit(500); // crawl more pages from now on
await robot.delete();

See Robot Management to run, schedule and manage robots.