A production-grade multi-agent framework for autonomous web exploration and real-time data extraction.
Live Demo β’ Architecture β’ Agents β’ ToT Reasoning
AWE is a sophisticated multi-agent system that autonomously explores websites to extract structured data. Unlike simple scrapers, AWE uses Tree of Thought (ToT) reasoning to plan, execute, and validate extraction strategies, enabling Small Language Models (SLMs) like llama-3.1-8b to perform on par with much larger models.
- Multi-Agent Coordination: 6 specialized agents working in concert.
- Cognitive Architecture: Uses ToT to "think" before acting.
- SLM Optimization: Optimized for fast inference on modest hardware.
- Live Extraction: Real-time fetching from any public URL.
AWE uses a Hub-and-Spoke architecture where the Orchestrator coordinates specialized agents that share a central state.
graph TD
User[User / Frontend] -->|Goal & URL| API[FastAPI Server]
API -->|Task| Orch[Orchestrator]
subgraph "AWE Core Engine"
Orch -->|Coordinates| Pool[Agent Pool]
Orch -->|Updates| State[Shared State]
subgraph "Reasoning Layer"
ToT[ToT Engine] -->|Empowers| Planner
ToT -->|Empowers| Extractor
end
subgraph "Agent Pool"
Observer[π Observer Agent]
Planner[π§ Planner Agent]
Executor[β‘ Executor Agent]
Extractor[π¦ Extractor Agent]
Validator[β
Validator Agent]
Learner[π Learner Agent]
end
Observer -->|Reads| Browser[Playwright Browser]
Executor -->|Drives| Browser
end
Extractor -->|Data| Results[Structured Data]
The system is composed of 6 specialized agents, each with a distinct role in the extraction pipeline:
- Role: The "Eyes" of the system.
- Responsibility: Analyzes the rendered DOM, takes screenshots, identifies page types (Listing, Detail, Home), and detects dynamic content (AJAX, Infinite Scroll).
- Output:
PageObservationobject with visual and structural context.
- Role: The "Brain" / Strategist.
- Responsibility: Uses Tree of Thought to generate high-level exploration strategies. It weighs trade-offs between different approaches (e.g., "Pagination Crawl" vs. "Search Query").
- Output:
ExplorationStrategycontaining a prioritized list of actions.
- Role: The "Hands" of the system.
- Responsibility: Interacts with the browser via Playwright. Executes planned actions like clicking, scrolling, typing, and navigating. Handles retries and error recovery.
- Output:
ActionExecutionResult(success/failure, new HTML).
- Role: The Data Miner.
- Responsibility: Parses content to extract the specific fields requested by the user. Adapts to different layouts (Cards, Tables, JSON-LD) using LLM-guided selectors.
- Output: Raw JSON data.
- Role: The Quality Control.
- Responsibility: Checks extracted data against validity rules (schema compliance, null checks, data types). Filters out noise and hallucinated content.
- Output:
ValidationReportand cleaned data.
- Role: The Optimizer.
- Responsibility: Remembers successful extraction patterns for specific domains. Creates reusable templates to speed up future runs on the same website.
- Output:
ExtractionTemplatestored in Knowledge Graph.
AWE implements a Tree of Thought (ToT) reasoning engine to solve the problem of complex web navigation and extraction. This allows the system to:
- Generate Thoughts: Propose multiple, distinct extraction strategies (e.g., "Try CSS selectors", "Try looking for JSON variables", "Try API interception").
- Evaluate: Score each strategy based on feasibility, confidence, and value.
- Search: Use Beam Search or BFS to explore the most promising strategies.
- Reflect: If a strategy fails, the system "reflects" on why and updates its plan.
Why ToT? Standard LLM calls fail on complex sites. ToT allows Small Language Models (SLMs) like
llama-3.1-8bto achieve SOTA performance by breaking the problem down and validating intermediate steps.
The extraction process follows a strictly defined 6-Phase Lifecycle:
sequenceDiagram
participant User
participant Orch as Orchestrator
participant Agents as Agent Pool
participant Browser
User->>Orch: Start Task (URL + Goal)
rect rgb(30, 30, 30)
note right of Orch: PHASE 1: Observation
Orch->>Browser: Goto URL
Orch->>Agents: Observer.observe()
Agents-->>Orch: PageObservation
end
rect rgb(40, 40, 40)
note right of Orch: PHASE 2: Planning (ToT)
Orch->>Agents: Planner.plan()
Agents->>Agents: Generate & Evaluate Thoughts
Agents-->>Orch: ExplorationStrategy
end
rect rgb(50, 50, 50)
note right of Orch: PHASE 3: Discovery
Orch->>Agents: Executor.execute(Strategy)
Agents->>Browser: Interact / Scroll / Click
Agents-->>Orch: List[Item URLs]
end
loop For Each Item
rect rgb(60, 60, 60)
note right of Orch: PHASE 4: Extraction
Orch->>Browser: Navigate to Item
Orch->>Agents: Extractor.extract()
Agents-->>Orch: Raw Data
end
end
rect rgb(70, 70, 70)
note right of Orch: PHASE 5: Validation
Orch->>Agents: Validator.validate(Data)
Agents-->>Orch: Validated Data
end
rect rgb(80, 80, 80)
note right of Orch: PHASE 6: Learning
Orch->>Agents: Learner.learn(Patterns)
Agents-->>Orch: Saved Template
end
Orch->>User: Final Result
- Framework: Next.js 16 (App Router)
- UI Library: React 19
- Styling: TailwindCSS 4
- State: React Hooks
- Theme: Glassmorphism (Dark Mode)
- API Framework: FastAPI
- Language: Python 3.11+
- Browser Control: Playwright (Async)
- LLM Interface: Groq SDK / Ollama
- Networking: httpx
- Parsing: BeautifulSoup4, lxml
- Reasoning Model:
llama-3.3-70b-versatile(Groq LPU) - SLM Option:
llama-3.1-8b-instant(High speed) - Vision Support: Experimental (GPT-4o / Llava)
The live demo on the landing page is a turbo-charged version of the pipeline designed for speed.
- Mode: Real-time Interactive
- Engine:
tot_extractor.py(Simplified ToT) - Latency: 3-8 seconds
- Capabilities: Single-page extraction, multiple strategies
To run the full multi-agent exploration (multi-page, deep crawling), use the /explore API endpoint.
Set these in your .env file:
# AI Provider
MODEL_PROVIDER=groq
GROQ_API_KEY=your_key_here
MODEL_NAME=llama-3.3-70b-versatile
# ToT Settings
TOT_ENABLED=true
TOT_MAX_THOUGHTS=3
TOT_SEARCH_STRATEGY=beam
# Agent Settings
HEADLESS=true
MAX_PAGES=10[AWE] Agentic Web Explorer Autonomous. Intelligent. Adaptive.