BetaWeekly digests are a beta — we're trying something new. Feedback welcome.
Agents get richer data sources as Firecrawl pushes Alexandria everywhere
September 28 – October 4, 2026
Firecrawl opened its Alexandria data layer to people-and-company enrichment and brought it inside ChatGPT and Codex, while Crawlee hardened its browser pool against hanging page closes and task-context leaks.
Enrichment lands in the same account agents already trust
Firecrawl's Alexandria layer grew well past web scraping this week. Agents can now search for people and enrich companies through Apollo, FullEnrich, and Data Legion using the same Firecrawl account and credits they already have — no second API key, no separate billing relationship. That covers looking up decision makers by title, company domain, location, role, skills, and work history, pulling work emails and employment history, and enriching firmographics, funding rounds, headcount, technologies, job postings, and news.
The pricing model is worth noting for anyone building cost-sensitive pipelines: discovery is free, and every tool surfaces its execution price before the call runs. That keeps enrichment from becoming a silent line item in an agent loop, which is the usual failure mode when you bolt a third-party data vendor onto an existing stack.
Alexandria reaches ChatGPT and Codex
The bigger distribution story is Alexandria running inside ChatGPT and Codex via the Firecrawl plugin. Agents in those conversations can now search and scrape the live web with access to 100+ data providers and Firecrawl's specialized indexes, without leaving the chat. Same pricing shape as above — free discovery and inspection, execution billed at each tool's listed rate.
Firecrawl's internal evaluation claims agents using Alexandria scored 21% higher on answer quality than those leaning on built-in web tools, across 845 tasks with blind AI judging. Treat vendor-run benchmarks with the usual skepticism, but the practical point stands: if your agent already lives in ChatGPT or Codex, the retrieval path just got considerably wider without any integration work.
Crawlee's browser pool gets defensive
Crawlee shipped a maintenance release that reads like accumulated scar tissue from production scraping. The browser pool now survives a page close that never settles, which is the kind of hang that silently strands workers and slowly starves a crawler of capacity. The same release stops notify() from leaking a timed-out task's context into the pool loop — a subtler bug, but one that produces confusing cross-task behavior when it fires.
Neither change is glamorous. Both are the sort of fix that matters most on long runs where a single wedged browser slot compounds over hours. If you run Crawlee at scale and have ever watched throughput decay without an obvious cause, this is the update to pull in.
Where this leaves the stack
Firecrawl is clearly betting that the data layer, not the scraper, is the durable product — enrichment and live web access converging into one connection that follows agents into whatever surface they run in. Crawlee continues playing the opposite role: the workhorse you run yourself, quietly getting sturdier. Neither direction is wrong, and for most teams the two will sit in the same pipeline.