Skip to content

Repository files navigation

Overview

DataDive is an enterprise knowledge management platform that enables intelligent search and question answering across heterogeneous enterprise data. The system combines Retrieval-Augmented Generation (RAG), hybrid retrieval, structured query execution, reranking, and role-based access control to retrieve accurate, context-aware information from both structured and unstructured sources.

Unlike conventional RAG systems that rely solely on vector search over documents, DataDive automatically classifies incoming queries, routes them through dedicated retrieval pipelines, combines the results using weighted reranking, and generates grounded responses using Large Language Models.

Frontend Repo: Click Here


System Workflow

1. Data Ingestion

The ingestion pipeline accepts multiple enterprise data formats including:

  • PDF
  • DOCX
  • PPTX
  • Images
  • CSV
  • JSON
  • SQL Databases

During ingestion, the system:

  • extracts document content and metadata
  • generates embeddings for semantic retrieval
  • classifies data into structured and unstructured sources
  • builds searchable indexes
  • creates PDF table-of-contents navigation for section-level lookup

2. Query Classification

Every incoming query is analysed before retrieval.

The classifier determines:

  • Structured Query
  • Unstructured Query
  • Hybrid Query

This routing allows the system to select the most appropriate retrieval strategy instead of treating every request as a semantic search problem.


3. Retrieval Pipeline

Depending on the query type:

Structured Pipeline

  • Converts natural language into SQL
  • Executes queries against relational databases
  • Returns validated structured results

Unstructured Pipeline

  • Performs semantic retrieval over indexed enterprise documents
  • Retrieves relevant chunks using vector similarity

Hybrid Pipeline

  • Executes both structured and semantic retrieval in parallel
  • Merges the retrieved evidence

4. Reranking

Retrieved candidates are reranked using a weighted scoring strategy based on:

  • Semantic similarity
  • Keyword overlap
  • Source reliability

The highest-ranked context is then forwarded to the language model for response generation.


5. Response Generation

The selected context is supplied to the LLM to generate grounded answers.

Responses are streamed to the client using Server-Sent Events (SSE), while access permissions are enforced through Role-Based Access Control (RBAC).


Architecture

                    Enterprise Data
      (PDF | DOCX | Images | CSV | SQL | JSON)
                          │
                          ▼
                Document Ingestion Pipeline
                          │
      Metadata Extraction & Document Classification
                          │
          ┌───────────────┴────────────────┐
          ▼                                ▼
 Structured Sources              Unstructured Sources
          │                                │
 SQL Schema Index              Chunking & Embeddings
          │                                │
          └───────────────┬────────────────┘
                          ▼
                 Hybrid Retrieval Engine
                          │
                Query Classification
      (Structured | Unstructured | Hybrid)
                          │
          ┌───────────────┴────────────────┐
          ▼                                ▼
      SQL Execution               Semantic Retrieval
          └───────────────┬────────────────┘
                          ▼
                Weighted Reranking
                          │
                  Context Construction
                          │
                          ▼
                   LLM Response
                          │
          RBAC + Streaming Response (SSE)

Key Features

  • Enterprise Retrieval-Augmented Generation (RAG)
  • Hybrid retrieval across structured and unstructured data
  • Automatic query classification
  • Natural language to SQL execution
  • Weighted reranking pipeline
  • Semantic search with vector embeddings
  • Metadata extraction during document ingestion
  • PDF table-of-contents based document navigation
  • Interactive knowledge visualizations
  • Role-Based Access Control (RBAC)
  • Real-time streaming responses using SSE
  • Blockchain-backed audit logging (research architecture)

Research

DataDive is the implementation of our research on enterprise knowledge retrieval.

Publications

  • ACT2025 (Scopus Indexed): Hybrid Retrieval System for Enterprise Knowledge Management

The proposed architecture achieved:

Metric Result
Top-5 Retrieval Accuracy 87.5%
SQL Query Correctness 92.1%

About

DataDive is a secure, scalable RAG system for enterprise knowledge management. It integrates multimodal data, provides AI-driven, context-aware answers, and ensures dynamic access control using ML. With blockchain-based audit trails and an interactive UI, DataDive enhances decision-making and data security.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages