Spaces:
Sleeping
Sleeping
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,35 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Multimodal AI Doctor – An Agentic AI Project
|
| 2 |
|
| 3 |
-
**Multimodal AI Doctor** is an **agentic multimodal assistant** built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
|
| 4 |
-
It combines **speech, vision, and reasoning** through a series of cooperating LLMs, simulating how a real doctor listens, observes, and responds concisely.
|
| 5 |
-
The system integrates **voice input, image analysis, clinical reasoning, and voice output** into a single pipeline.
|
| 6 |
|
| 7 |
---
|
| 8 |
|
| 9 |
## Features
|
| 10 |
|
| 11 |
-
* Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
|
| 12 |
-
* Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
|
| 13 |
-
* Generate a concise medical-style response (2 sentences maximum, human-like tone)
|
| 14 |
-
* Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
|
| 15 |
-
* Gradio-based interactive UI
|
| 16 |
|
| 17 |
---
|
| 18 |
|
| 19 |
## Project Structure
|
| 20 |
|
| 21 |
```
|
|
|
|
| 22 |
.
|
| 23 |
├── app.py # Gradio UI + main workflow
|
| 24 |
-
├──
|
| 25 |
-
├──
|
| 26 |
-
├──
|
| 27 |
├── requirements.txt # Python dependencies
|
| 28 |
├── .env # Environment variables (API keys)
|
| 29 |
-
├── .gitignore # Ignore venv,
|
| 30 |
├── images/ # Folder for saving test/sample images
|
| 31 |
└── README.md # Documentation
|
| 32 |
-
|
|
|
|
| 33 |
|
| 34 |
---
|
| 35 |
|
|
@@ -37,22 +50,22 @@ The system integrates **voice input, image analysis, clinical reasoning, and voi
|
|
| 37 |
|
| 38 |
The system uses **multiple LLM agents** to process multimodal input step by step:
|
| 39 |
|
| 40 |
-
1. **Symptom Agent** – extracts structured meaning from patient speech (via Whisper transcription).
|
| 41 |
-
2. **Vision Agent** – analyzes uploaded medical images (X-ray, MRI, scan).
|
| 42 |
-
3. **Reasoning Agent** – integrates speech and image findings into a medical interpretation.
|
| 43 |
-
4. **Response Agent** – formats the answer in a concise, empathetic, doctor-style tone (≤ 2 sentences).
|
| 44 |
-
5. **Voice Agent** – delivers the response using ElevenLabs (WAV, fallback gTTS).
|
| 45 |
|
| 46 |
-
This makes the project an **agentic AI pipeline**, where multiple specialized models cooperate to mimic a doctor’s diagnostic process.
|
| 47 |
|
| 48 |
---
|
| 49 |
|
| 50 |
## Requirements
|
| 51 |
|
| 52 |
-
* Python 3.10 or higher
|
| 53 |
-
* FFmpeg installed and available in PATH (required by pydub)
|
| 54 |
-
* A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
|
| 55 |
-
* An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
|
| 56 |
|
| 57 |
---
|
| 58 |
|
|
@@ -63,7 +76,7 @@ This makes the project an **agentic AI pipeline**, where multiple specialized mo
|
|
| 63 |
```bash
|
| 64 |
git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
|
| 65 |
cd ai-doctor-2.0-voice-and-vision
|
| 66 |
-
|
| 67 |
|
| 68 |
2. Create and activate a virtual environment:
|
| 69 |
|
|
@@ -99,7 +112,7 @@ This makes the project an **agentic AI pipeline**, where multiple specialized mo
|
|
| 99 |
Start the Gradio app:
|
| 100 |
|
| 101 |
```bash
|
| 102 |
-
python
|
| 103 |
```
|
| 104 |
|
| 105 |
The app will launch locally at:
|
|
@@ -158,4 +171,6 @@ For questions, issues, or collaboration, please contact:
|
|
| 158 |
**Email:** [sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)
|
| 159 |
**LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
|
| 160 |
|
| 161 |
-
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: Multimodal AI Doctor – An Agentic AI Project
|
| 3 |
+
emoji: 🩺
|
| 4 |
+
colorFrom: indigo
|
| 5 |
+
colorTo: blue
|
| 6 |
+
sdk: gradio
|
| 7 |
+
sdk_version: 3.50.2
|
| 8 |
+
app_file: app.py
|
| 9 |
+
pinned: false
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
# Multimodal AI Doctor – An Agentic AI Project
|
| 13 |
|
| 14 |
+
**Multimodal AI Doctor** is an **agentic multimodal assistant** built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
|
| 15 |
+
It combines **speech, vision, and reasoning** through a series of cooperating LLMs, simulating how a real doctor listens, observes, and responds concisely.
|
| 16 |
+
The system integrates **voice input, image analysis, clinical reasoning, and voice output** into a single pipeline.
|
| 17 |
|
| 18 |
---
|
| 19 |
|
| 20 |
## Features
|
| 21 |
|
| 22 |
+
* Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
|
| 23 |
+
* Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
|
| 24 |
+
* Generate a concise medical-style response (2 sentences maximum, human-like tone)
|
| 25 |
+
* Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
|
| 26 |
+
* Gradio-based interactive UI
|
| 27 |
|
| 28 |
---
|
| 29 |
|
| 30 |
## Project Structure
|
| 31 |
|
| 32 |
```
|
| 33 |
+
|
| 34 |
.
|
| 35 |
├── app.py # Gradio UI + main workflow
|
| 36 |
+
├── brain\_of\_the\_doctor.py # Image encoding + Groq multimodal analysis
|
| 37 |
+
├── voice\_of\_the\_patient.py # Audio recording + Groq Whisper transcription
|
| 38 |
+
├── voice\_of\_the\_doctor.py # ElevenLabs + gTTS text-to-speech
|
| 39 |
├── requirements.txt # Python dependencies
|
| 40 |
├── .env # Environment variables (API keys)
|
| 41 |
+
├── .gitignore # Ignore venv, **pycache**, .env, etc.
|
| 42 |
├── images/ # Folder for saving test/sample images
|
| 43 |
└── README.md # Documentation
|
| 44 |
+
|
| 45 |
+
````
|
| 46 |
|
| 47 |
---
|
| 48 |
|
|
|
|
| 50 |
|
| 51 |
The system uses **multiple LLM agents** to process multimodal input step by step:
|
| 52 |
|
| 53 |
+
1. **Symptom Agent** – extracts structured meaning from patient speech (via Whisper transcription).
|
| 54 |
+
2. **Vision Agent** – analyzes uploaded medical images (X-ray, MRI, scan).
|
| 55 |
+
3. **Reasoning Agent** – integrates speech and image findings into a medical interpretation.
|
| 56 |
+
4. **Response Agent** – formats the answer in a concise, empathetic, doctor-style tone (≤ 2 sentences).
|
| 57 |
+
5. **Voice Agent** – delivers the response using ElevenLabs (WAV, fallback gTTS).
|
| 58 |
|
| 59 |
+
This makes the project an **agentic AI pipeline**, where multiple specialized models cooperate to mimic a doctor’s diagnostic process.
|
| 60 |
|
| 61 |
---
|
| 62 |
|
| 63 |
## Requirements
|
| 64 |
|
| 65 |
+
* Python 3.10 or higher
|
| 66 |
+
* FFmpeg installed and available in PATH (required by pydub)
|
| 67 |
+
* A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
|
| 68 |
+
* An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
|
| 69 |
|
| 70 |
---
|
| 71 |
|
|
|
|
| 76 |
```bash
|
| 77 |
git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
|
| 78 |
cd ai-doctor-2.0-voice-and-vision
|
| 79 |
+
````
|
| 80 |
|
| 81 |
2. Create and activate a virtual environment:
|
| 82 |
|
|
|
|
| 112 |
Start the Gradio app:
|
| 113 |
|
| 114 |
```bash
|
| 115 |
+
python app.py
|
| 116 |
```
|
| 117 |
|
| 118 |
The app will launch locally at:
|
|
|
|
| 171 |
**Email:** [sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)
|
| 172 |
**LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
|
| 173 |
|
| 174 |
+
```
|
| 175 |
+
|
| 176 |
+
---
|