Cloudflare has changed how its Markdown for Agents feature performs HTML-to-Markdown conversion, according to a changelog entry on the company's developer documentation site. The conversion now runs through an in-process streaming engine at the edge, rather than buffering the HTML response and dispatching it to a separate conversion service. Cloudflare states that this approach processes content as it arrives and reduces both conversion overhead and memory use.
The same changelog entry documents a second change: the conversion limit has been raised. Markdown for Agents now accepts up to 6 MiB (6,291,456 bytes) of decompressed HTML, up from the previous 2 MiB (2,097,152 bytes). Cloudflare specifies that the ceiling applies after decompression, not to the compressed size of the response, which matters for anyone reasoning about how large a page can be before conversion stops.
Two response-header behaviours also changed. Converted responses no longer emit the x-markdown-tokens or x-original-tokens headers, so clients that relied on those values must now compute token counts themselves. Separately, Content-Length is removed from converted responses rather than recalculated, which Cloudflare attributes to the Markdown body being streamed.
For developers who build tooling on top of these responses, the header changes are the most likely source of breakage. A pipeline that reads x-markdown-tokens to budget context, log usage, or decide whether to truncate will silently lose that input and needs its own tokenizer or counting step. The removal of Content-Length is a related consequence of streaming: a response whose length is not known in advance cannot advertise a fixed length, so consumers that pre-allocate buffers or display progress based on that header need a different strategy.
The architecture change is the more structural of the two. Buffering an HTML response and forwarding it to a separate conversion service implies at least one extra hop and a full copy of the document held in memory before conversion begins. An in-process streaming engine at the edge, as described, converts as bytes arrive, which is consistent with the stated reductions in overhead and memory use. Cloudflare does not publish latency or throughput figures in the supplied text, so the size of any performance gain is not quantified here.
The raised limit and the streaming design are complementary but distinct. A larger accepted input means more documents qualify for conversion; streaming means the conversion path does not need the whole document resident before producing output. Cloudflare's note that the 6 MiB figure is measured after decompression is a practical detail: a highly compressible page can be well under the limit on the wire while exceeding it once expanded, and the decompressed figure is the one that governs.
The changelog points readers to the Markdown for Agents documentation for further information, but the supplied material does not include that documentation, pricing details, regional availability, or any statement about which plans or accounts receive the feature. It also does not describe how the previous separate conversion service behaved in detail, nor whether the in-process engine produces byte-identical Markdown output to the prior path. Those are open questions rather than documented facts.
For freelancers and small studios maintaining scraping, archiving, or AI-ingestion pipelines, the practical reading is that this is a behaviour change to code that consumes converted responses, not merely an internal optimisation. The safest response is to audit any dependency on the two removed token headers and on Content-Length, and to add explicit token counting where those values were previously trusted. That work is small but easy to overlook because the endpoint's purpose is unchanged.
There is also a capacity consideration. Teams that previously skipped or chunked pages above 2 MiB of decompressed HTML may now be able to send them directly, up to the new 6 MiB ceiling. Whether that is desirable depends on downstream token budgets, since a larger converted document consumes more context regardless of how efficiently it was produced. The limit change expands what is accepted, not what is advisable to feed into a model.
Editorially, the significance of this release is modest but concrete: a documented change to conversion mechanics, a documented limit increase, and two documented header removals, all attributable to Cloudflare's own changelog. It is not evidence of a broader platform shift, and no third-party measurements are supplied to confirm the claimed efficiency gains. Treat the overhead and memory statements as vendor claims about the new design rather than independently verified results.
What remains unknown is how the streaming engine handles malformed or truncated HTML, whether conversion output differs from the previous service in edge cases, and whether the removed headers will be replaced by an alternative mechanism. Cloudflare's text does not address these points. Developers who depend on precise token accounting should assume they now own that responsibility until documentation says otherwise.
The bottom line for this audience: update consumers of Markdown for Agents responses to stop expecting x-markdown-tokens, x-original-tokens, and Content-Length; recalculate token counts locally; and revisit any size-based gating that assumed a 2 MiB decompressed ceiling. The conversion itself is described as faster and lighter, but the only changes that require action are the limit and the headers.