ALLAM-RAG: A Context-Aware Saudi Arabic Conversational AI System for Alzheimer’s Support

Overview

ALLAM-RAG is a Saudi Arabic conversational AI system that combines Retrieval-Augmented Generation (RAG) with the ALLAM Large Language Model to provide context-aware assistance for Alzheimer's patients and caregivers.

The system retrieves relevant information from specialized knowledge bases using hybrid vector search and reranking before generating accurate, personalized responses through the ALLAM language model. By combining retrieval with language generation, the system minimizes hallucinations while delivering reliable and natural Arabic conversations tailored to the Saudi dialect.


Features

  • Saudi Arabic conversational assistant
  • Retrieval-Augmented Generation (RAG) architecture
  • Powered by the ALLAM Large Language Model
  • Hybrid semantic and keyword retrieval using Weaviate
  • Multilingual reranking using Cohere Rerank
  • Context-aware response generation
  • Alzheimer's patient support
  • General Saudi knowledge support
  • Real-time time and day awareness
  • Arabic text normalization and intent detection

System Architecture

User Query
      │
      ▼
Query Classification
      │
      ▼
Knowledge Base Selection
      │
      ▼
Weaviate Hybrid Search
      │
      ▼
Cohere Reranker
      │
      ▼
Relevant Context
      │
      ▼
ALLAM Language Model
      │
      ▼
Generated Response

Technologies Used

  • Python
  • ALLAM-7B-Instruct
  • Transformers
  • Hugging Face
  • Weaviate Vector Database
  • Cohere Rerank API
  • NumPy

Knowledge Bases

The system utilizes two specialized knowledge bases:

Alzheimer's Knowledge Base

Contains information related to:

  • Patient identity
  • Memory assistance
  • Medication reminders
  • Daily activities
  • Alzheimer's symptoms
  • Caregiving guidance
  • Orientation support

General Knowledge Base

Contains Saudi Arabic information including:

  • Islamic knowledge
  • Saudi Arabia
  • Culture
  • Geography
  • Public information

Retrieval Pipeline

The retrieval pipeline consists of:

  1. Query preprocessing and normalization
  2. Query classification
  3. Knowledge base selection
  4. Hybrid retrieval from Weaviate
  5. Cohere multilingual reranking
  6. Context selection
  7. Response generation using ALLAM

Example Queries

Alzheimer's Support

  • مين أنا؟
  • وين ساكن؟
  • كيف حالتي الصحية؟
  • نسيت أخذ الدواء
  • ليه أنا هنا؟

Project Structure

app.py
chatbot.py
rag_pipeline.py
weaviate_server.py
requirements.txt
README.md

Running Locally

Install dependencies:

pip install -r requirements.txt

Configure the required environment variables:

  • COHERE_API_KEY
  • HF_TOKEN
  • WEAVIATE_URL (if applicable)
  • WEAVIATE_API_KEY (if applicable)

Launch the application:

python app.py

System Architecture

System Architecture

Deployment

The project is designed for deployment using:

  • Hugging Face Models
  • ALLAM language model
  • Weaviate Vector Database
  • Cohere Language model

Demo

Responses Demo

Future Improvements

  • Multi-turn conversation memory
  • Expanded healthcare knowledge base

Author

Shahad Aljohani

Recent Computer Science Undergraduate


Privacy Notice

This repository contains a public research version. Private implementation details, API keys, deployment configurations, and some internal pipeline components are excluded.

The complete implementation is maintained privately.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support