> ## Documentation Index
> Fetch the complete documentation index at: https://help.ciarem.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Website sources: index your whole site

> How Ciarem finds the pages of a website, how you choose which ones to index, and how it keeps them up to date with a weekly recrawl.

A website source is your whole site, not a single page. You add the address. Ciarem finds the pages behind it. You choose which pages your agents should learn from. Every week, Ciarem checks those pages for changes and tells you when the site has new pages.

## How Ciarem finds your pages

Discovery runs while you wait in the dialog. It takes about ten seconds at most. Ciarem tries the most complete source first:

1. **Your sitemap.** Ciarem reads `robots.txt` to find the sitemap it declares, then tries `/sitemap.xml` and `/wp-sitemap.xml`. Most sites (WordPress, Webflow, Shopify, Next.js and others) list every public page there, so this is the usual result.
2. **Links on your pages.** If there is no sitemap, Ciarem reads the links on the address you entered. Then it reads the links on the pages that address points to. It only follows links on the same site.
3. **Common pages.** Ciarem also checks addresses most sites have, such as `/about`, `/pricing`, `/faq`, `/contact`, `/blog`, `/terms` and `/privacy`. It keeps the ones that really exist.
4. **The address you typed.** If nothing else turns up, Ciarem still indexes that page on its own.

Keep in mind:

* `www.example.com` and `example.com` count as the same site. If the one you typed does not respond, Ciarem tries the other one.
* Ciarem lists up to **100 pages** per site. It skips cart, checkout and login pages, admin areas, and files such as PDFs and images. To index a PDF, add it as a [file source](/ai-agent/knowledge-bases).
* Language versions of a page (for example `/es/` and `/pt/`) are listed too. You decide which ones to keep.

## Add a website

<Steps>
  <Step title="Open a knowledge base">
    Click **Add information**, then **Website**.
  </Step>

  <Step title="Name the source and type the address">
    Give the source a name and enter the address of the site.

    <img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-dialog.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=199515faf2363b0ade3d9e4f2b3b3714" alt="The website dialog: a name, the site address, and the Find pages button" width="1680" height="1114" data-path="images/ciarem-help-website-dialog.png" />
  </Step>

  <Step title="Click Find pages">
    The pages appear grouped like folders. The address you typed comes first, then groups such as `blog`, `es` or `pt`, then single pages. All pages start selected.

    <img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-pages.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=7564b199e7d1cc729b27db5c5864a6d8" alt="The pages Ciarem found on ciarem.ai, grouped by section, all selected" width="1680" height="1114" data-path="images/ciarem-help-website-pages.png" />
  </Step>

  <Step title="Choose the pages to index">
    * The checkbox of a group selects or clears every page inside it. One click covers the whole blog, or only its Portuguese posts. The checkbox shows a dash when only some pages in the group are selected.
    * Click the arrow next to a group to open it and pick individual pages.
    * Use the filter box to find a page by name or address.

          <img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-pages-groups.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=3f5e852a54a4218c9e03429c940599f9" alt="The blog group opened, with its Portuguese posts unselected and the group checkbox showing a partial selection" width="1680" height="1114" data-path="images/ciarem-help-website-pages-groups.png" />
  </Step>

  <Step title="Click Add N pages">
    The counter shows how many pages you are about to add. Click **Add N pages** to finish.
  </Step>
</Steps>

<Tip>
  Unselect the pages your agent should not learn from, such as blog archives, tag pages, or languages you do not serve. A smaller set of good pages gives better answers than the whole site.
</Tip>

## What happens after you add it

Ciarem fetches and indexes each page separately. The source card shows the number of pages and fragments. To see every page and its state, open **⋯ → See pages**:

* **Indexed**: the page is part of the knowledge base.
* **Indexing**: the page is still being processed.
* **Failed**: the page could not be indexed. The list shows why. For example: the page could not be reached, the site answered with an error, the page is not HTML, it has no readable text, it needs JavaScript to show its content, or the site blocked the request.

<img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-page-status.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=9248cc711b01dbea25178175fcd0ff0f" alt="The page list of a website source: every page with its title, address, state and number of fragments" width="1680" height="1114" data-path="images/ciarem-help-website-page-status.png" />

A few failed pages do not fail the whole source. The source stays ready, and the card tells you how many pages failed, so you can decide whether they matter.

Ciarem leaves out text that repeats on every page, such as menus, footers and cookie banners. Sites built with JavaScript are supported. When the content of a page only appears after JavaScript runs, Ciarem renders the page before indexing it.

## Keeping it up to date: the weekly recrawl

Once a week, Ciarem visits every page of the source again:

* **Pages that did not change** stay as they are. Ciarem compares the text of the page with the version it indexed, so it only re-indexes a page when the content actually changed.
* **Pages that changed** are re-indexed one by one. The other pages are not touched.
* **Pages that fail** keep the content they had. They are marked as failed in the page list. A temporary outage never removes answers your agent already gives.

New pages on the site are **not added automatically**. The card tells you how many new pages Ciarem found. Click **Review** to see them.

<img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-card.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=1156128f5ecccb0b8b7a9e463539ef37" alt="A website source after a recrawl: the card shows how many new pages were found and a Review button" width="1680" height="1114" data-path="images/ciarem-help-website-card.png" />

The review dialog groups the new pages the same way as the page picker. Select the pages your agent should learn from and click **Add N pages**. Ciarem indexes the source again with the new pages. Click **Not now** to skip them. Skipped pages appear again in a later review if the site still has them.

<img src="https://mintcdn.com/ciaremaai/4qDbKW3sTNK9-WSj/images/ciarem-help-website-new-pages.png?fit=max&auto=format&n=4qDbKW3sTNK9-WSj&q=85&s=472c36b0c15551ebb2e03df7cf12cce5" alt="Reviewing the new pages found on the site, grouped by section, before adding them" width="1680" height="1114" data-path="images/ciarem-help-website-new-pages.png" />

## Good to know

* Website sources created before this feature keep their single page. To index the whole site, add it again as a new website source and delete the old one.
* Discovery only reads public pages. Pages behind a login or a paywall are not found.
* The weekly recrawl runs on its own. There is nothing to schedule.
