Raw HTML makes poor context for an LLM. A typical page is mostly navigation, scripts, cookie banners and footer links. Putting all of that into an embedding model or a prompt wastes tokens and lowers retrieval quality. What you want is the content : headings, paragraphs, lists and links, as markdown, plus enough metadata to cite the source and to tell when it goes stale. This guide uses AgentSearch Web Extract ( agentsearch-web-extract-v1 ) to turn a URL into RAG-ready markdown. It covers ...