What can the AI train on?

The four source types, and what each is good for.

C
ChatteringJuly 22, 2026

Chattering answers only from content you give it. There are four source types.

Website

Give it a URL and it crawls the pages it can reach, splitting each into passages. Best for public docs, marketing pages and existing help centres.

Pages behind a login cannot be crawled — paste that content as text instead.

The crawler follows links from the starting URL and picks up pages on the same domain. It skips PDFs and other binary files. If your site has an XML sitemap, Chattering uses it to find pages it might otherwise miss.

Recrawling is manual. If you update your site, go to the source row in your agent's Sources tab and click Recrawl to pick up the changes.

Text

Paste anything: an internal policy, a pricing explanation, a support macro you already use. Best for knowledge that lives in someone's head or in a document rather than on a page.

There is no hard character limit on a text source. Very large pastes take longer to index, so if you have a large document, consider splitting it into logical sections — one source per product area, for example — so you can update each part independently later.

Q&A

One question, one exact answer. Use this when the wording of the answer matters — refund policy, security posture, anything with a phrase you need said precisely.

Q&A pairs are the strongest way to steer a specific answer. Unlike crawled pages, the answer is reproduced closely rather than paraphrased from a passage, so there is no risk of the AI rewording something you need kept exact. A good Q&A pair anticipates how the customer will phrase the question, not just the ideal answer.

File

Upload a plain-text file. PDF and Word are not supported — the uploader says so rather than accepting a file it cannot read. Copy the text out and paste it as a text source instead.

Plain-text files (.txt) are the only supported format. The maximum file size is 5 MB. If your file is larger, split it into sections before uploading.

Your own help centre

Articles you publish in Chattering's help centre can be synced back into the agent's knowledge in one click, from AI agents → Content → Help articles. Writing an article therefore improves the AI at the same time.

When you publish a new article, it does not sync automatically. Open the Help articles panel and click Sync to push the latest articles into the index.

How the AI uses your content

Chattering does not retrain or fine-tune the underlying language model. Instead, it indexes your content into a separate search layer. When a question comes in, Chattering retrieves the most relevant passages from that layer and passes them to the model as context. The model then composes an answer grounded in those passages.

This means the base model's general knowledge of language and reasoning does not change when you add sources. Only the retrieval layer changes. Think of it less as "training" and more as giving the AI a set of reference documents to read before it answers each question.

Data privacy

Your content is stored within your workspace only. Chattering does not share your indexed material with other workspaces, and it does not use your content to improve the base model or any shared system. Each workspace's knowledge index is isolated from every other.

If you delete a source, the passages derived from it are removed from the index. Conversation transcripts and indexed content are subject to the data retention settings in your workspace.

What the AI will not do

The AI will not draw on general internet knowledge when answering customer questions — it only answers from what you have indexed. If a question falls outside your indexed content, the conversation is handed off to a human agent rather than guessing.

The AI does not learn from individual conversations automatically. What customers type in chat does not update the index. You control the content; the AI reads it.

Tips for better results

Start by looking at what your customers already ask. The support conversations you have now are the best guide to which content gaps to fill first.

For recurring questions about prices, policies, or step-by-step procedures, Q&A pairs usually work better than crawled pages. The answer is reproduced closely, so there is no risk of the AI paraphrasing something you need kept exact.

Keep individual text sources focused on one topic. A 300-word source covering one feature typically performs better than a 3,000-word document covering ten. The retrieval step matches questions to passages — specific content is easier to retrieve accurately than a dense mix of topics.

Update sources when your content changes. The index does not refresh automatically when you edit your website. If you pasted content by hand, edit the source directly in the Sources tab. If you added a URL, click Recrawl after making significant updates to that site.

Need more help? Browse all articles or ask the assistant in the chat — it answers from these docs and can connect you with the team.

Was this helpful?

Related articles

What can the AI train on? | Chattering Help Center | Chattering.ai