DataDive is an enterprise knowledge management platform that enables intelligent search and question answering across heterogeneous enterprise data. The system combines Retrieval-Augmented Generation (RAG), hybrid retrieval, structured query execution, reranking, and role-based access control to retrieve accurate, context-aware information from both structured and unstructured sources.
Unlike conventional RAG systems that rely solely on vector search over documents, DataDive automatically classifies incoming queries, routes them through dedicated retrieval pipelines, combines the results using weighted reranking, and generates grounded responses using Large Language Models.
Frontend Repo: Click Here
The ingestion pipeline accepts multiple enterprise data formats including:
- DOCX
- PPTX
- Images
- CSV
- JSON
- SQL Databases
During ingestion, the system:
- extracts document content and metadata
- generates embeddings for semantic retrieval
- classifies data into structured and unstructured sources
- builds searchable indexes
- creates PDF table-of-contents navigation for section-level lookup
Every incoming query is analysed before retrieval.
The classifier determines:
- Structured Query
- Unstructured Query
- Hybrid Query
This routing allows the system to select the most appropriate retrieval strategy instead of treating every request as a semantic search problem.
Depending on the query type:
Structured Pipeline
- Converts natural language into SQL
- Executes queries against relational databases
- Returns validated structured results
Unstructured Pipeline
- Performs semantic retrieval over indexed enterprise documents
- Retrieves relevant chunks using vector similarity
Hybrid Pipeline
- Executes both structured and semantic retrieval in parallel
- Merges the retrieved evidence
Retrieved candidates are reranked using a weighted scoring strategy based on:
- Semantic similarity
- Keyword overlap
- Source reliability
The highest-ranked context is then forwarded to the language model for response generation.
The selected context is supplied to the LLM to generate grounded answers.
Responses are streamed to the client using Server-Sent Events (SSE), while access permissions are enforced through Role-Based Access Control (RBAC).
Enterprise Data
(PDF | DOCX | Images | CSV | SQL | JSON)
│
▼
Document Ingestion Pipeline
│
Metadata Extraction & Document Classification
│
┌───────────────┴────────────────┐
▼ ▼
Structured Sources Unstructured Sources
│ │
SQL Schema Index Chunking & Embeddings
│ │
└───────────────┬────────────────┘
▼
Hybrid Retrieval Engine
│
Query Classification
(Structured | Unstructured | Hybrid)
│
┌───────────────┴────────────────┐
▼ ▼
SQL Execution Semantic Retrieval
└───────────────┬────────────────┘
▼
Weighted Reranking
│
Context Construction
│
▼
LLM Response
│
RBAC + Streaming Response (SSE)
- Enterprise Retrieval-Augmented Generation (RAG)
- Hybrid retrieval across structured and unstructured data
- Automatic query classification
- Natural language to SQL execution
- Weighted reranking pipeline
- Semantic search with vector embeddings
- Metadata extraction during document ingestion
- PDF table-of-contents based document navigation
- Interactive knowledge visualizations
- Role-Based Access Control (RBAC)
- Real-time streaming responses using SSE
- Blockchain-backed audit logging (research architecture)
DataDive is the implementation of our research on enterprise knowledge retrieval.
Publications
- ACT2025 (Scopus Indexed): Hybrid Retrieval System for Enterprise Knowledge Management
The proposed architecture achieved:
| Metric | Result |
|---|---|
| Top-5 Retrieval Accuracy | 87.5% |
| SQL Query Correctness | 92.1% |