Skip to content

Latest commit

Β 

History

11 Commits

Folders and files

Repository files navigation

AWE - Agentic Web Explorer Cognexus

AWE Logo

A production-grade multi-agent framework for autonomous web exploration and real-time data extraction.

Next.js FastAPI Groq Playwright Tree of Thought

Live Demo β€’ Architecture β€’ Agents β€’ ToT Reasoning


🎯 Overview

AWE is a sophisticated multi-agent system that autonomously explores websites to extract structured data. Unlike simple scrapers, AWE uses Tree of Thought (ToT) reasoning to plan, execute, and validate extraction strategies, enabling Small Language Models (SLMs) like llama-3.1-8b to perform on par with much larger models.

Key Capabilities

  • Multi-Agent Coordination: 6 specialized agents working in concert.
  • Cognitive Architecture: Uses ToT to "think" before acting.
  • SLM Optimization: Optimized for fast inference on modest hardware.
  • Live Extraction: Real-time fetching from any public URL.

πŸ—οΈ Architecture

AWE uses a Hub-and-Spoke architecture where the Orchestrator coordinates specialized agents that share a central state.

graph TD
    User[User / Frontend] -->|Goal & URL| API[FastAPI Server]
    API -->|Task| Orch[Orchestrator]
    
    subgraph "AWE Core Engine"
        Orch -->|Coordinates| Pool[Agent Pool]
        Orch -->|Updates| State[Shared State]
        
        subgraph "Reasoning Layer"
            ToT[ToT Engine] -->|Empowers| Planner
            ToT -->|Empowers| Extractor
        end
        
        subgraph "Agent Pool"
            Observer[πŸ‘€ Observer Agent]
            Planner[🧠 Planner Agent]
            Executor[⚑ Executor Agent]
            Extractor[πŸ“¦ Extractor Agent]
            Validator[βœ… Validator Agent]
            Learner[πŸ“š Learner Agent]
        end
        
        Observer -->|Reads| Browser[Playwright Browser]
        Executor -->|Drives| Browser
    end
    
    Extractor -->|Data| Results[Structured Data]
Loading

πŸ€– Multi-Agent System

The system is composed of 6 specialized agents, each with a distinct role in the extraction pipeline:

1. πŸ‘€ Observer Agent

  • Role: The "Eyes" of the system.
  • Responsibility: Analyzes the rendered DOM, takes screenshots, identifies page types (Listing, Detail, Home), and detects dynamic content (AJAX, Infinite Scroll).
  • Output: PageObservation object with visual and structural context.

2. 🧠 Planner Agent

  • Role: The "Brain" / Strategist.
  • Responsibility: Uses Tree of Thought to generate high-level exploration strategies. It weighs trade-offs between different approaches (e.g., "Pagination Crawl" vs. "Search Query").
  • Output: ExplorationStrategy containing a prioritized list of actions.

3. ⚑ Executor Agent

  • Role: The "Hands" of the system.
  • Responsibility: Interacts with the browser via Playwright. Executes planned actions like clicking, scrolling, typing, and navigating. Handles retries and error recovery.
  • Output: ActionExecutionResult (success/failure, new HTML).

4. πŸ“¦ Extractor Agent

  • Role: The Data Miner.
  • Responsibility: Parses content to extract the specific fields requested by the user. Adapts to different layouts (Cards, Tables, JSON-LD) using LLM-guided selectors.
  • Output: Raw JSON data.

5. βœ… Validator Agent

  • Role: The Quality Control.
  • Responsibility: Checks extracted data against validity rules (schema compliance, null checks, data types). Filters out noise and hallucinated content.
  • Output: ValidationReport and cleaned data.

6. πŸ“š Learner Agent

  • Role: The Optimizer.
  • Responsibility: Remembers successful extraction patterns for specific domains. Creates reusable templates to speed up future runs on the same website.
  • Output: ExtractionTemplate stored in Knowledge Graph.

🌳 Tree of Thought Engine

AWE implements a Tree of Thought (ToT) reasoning engine to solve the problem of complex web navigation and extraction. This allows the system to:

  1. Generate Thoughts: Propose multiple, distinct extraction strategies (e.g., "Try CSS selectors", "Try looking for JSON variables", "Try API interception").
  2. Evaluate: Score each strategy based on feasibility, confidence, and value.
  3. Search: Use Beam Search or BFS to explore the most promising strategies.
  4. Reflect: If a strategy fails, the system "reflects" on why and updates its plan.

Why ToT? Standard LLM calls fail on complex sites. ToT allows Small Language Models (SLMs) like llama-3.1-8b to achieve SOTA performance by breaking the problem down and validating intermediate steps.


πŸ”„ Workflow

The extraction process follows a strictly defined 6-Phase Lifecycle:

sequenceDiagram
    participant User
    participant Orch as Orchestrator
    participant Agents as Agent Pool
    participant Browser
    
    User->>Orch: Start Task (URL + Goal)
    
    rect rgb(30, 30, 30)
        note right of Orch: PHASE 1: Observation
        Orch->>Browser: Goto URL
        Orch->>Agents: Observer.observe()
        Agents-->>Orch: PageObservation
    end
    
    rect rgb(40, 40, 40)
        note right of Orch: PHASE 2: Planning (ToT)
        Orch->>Agents: Planner.plan()
        Agents->>Agents: Generate & Evaluate Thoughts
        Agents-->>Orch: ExplorationStrategy
    end
    
    rect rgb(50, 50, 50)
        note right of Orch: PHASE 3: Discovery
        Orch->>Agents: Executor.execute(Strategy)
        Agents->>Browser: Interact / Scroll / Click
        Agents-->>Orch: List[Item URLs]
    end
    
    loop For Each Item
        rect rgb(60, 60, 60)
            note right of Orch: PHASE 4: Extraction
            Orch->>Browser: Navigate to Item
            Orch->>Agents: Extractor.extract()
            Agents-->>Orch: Raw Data
        end
    end
    
    rect rgb(70, 70, 70)
        note right of Orch: PHASE 5: Validation
        Orch->>Agents: Validator.validate(Data)
        Agents-->>Orch: Validated Data
    end
    
    rect rgb(80, 80, 80)
        note right of Orch: PHASE 6: Learning
        Orch->>Agents: Learner.learn(Patterns)
        Agents-->>Orch: Saved Template
    end
    
    Orch->>User: Final Result
Loading

πŸ› οΈ Tech Stack

Frontend (awe-landing)

  • Framework: Next.js 16 (App Router)
  • UI Library: React 19
  • Styling: TailwindCSS 4
  • State: React Hooks
  • Theme: Glassmorphism (Dark Mode)

Backend (awe-agentic-web-explorer)

  • API Framework: FastAPI
  • Language: Python 3.11+
  • Browser Control: Playwright (Async)
  • LLM Interface: Groq SDK / Ollama
  • Networking: httpx
  • Parsing: BeautifulSoup4, lxml

AI Models & Engines

  • Reasoning Model: llama-3.3-70b-versatile (Groq LPU)
  • SLM Option: llama-3.1-8b-instant (High speed)
  • Vision Support: Experimental (GPT-4o / Llava)

πŸš€ Live Demo (/demo)

The live demo on the landing page is a turbo-charged version of the pipeline designed for speed.

  • Mode: Real-time Interactive
  • Engine: tot_extractor.py (Simplified ToT)
  • Latency: 3-8 seconds
  • Capabilities: Single-page extraction, multiple strategies

To run the full multi-agent exploration (multi-page, deep crawling), use the /explore API endpoint.


βš™οΈ Configuration

Set these in your .env file:

# AI Provider
MODEL_PROVIDER=groq
GROQ_API_KEY=your_key_here
MODEL_NAME=llama-3.3-70b-versatile

# ToT Settings
TOT_ENABLED=true
TOT_MAX_THOUGHTS=3
TOT_SEARCH_STRATEGY=beam

# Agent Settings
HEADLESS=true
MAX_PAGES=10

[AWE] Agentic Web Explorer Autonomous. Intelligent. Adaptive.

About

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages