Sayeem26s commited on
Commit
7074753
·
verified ·
1 Parent(s): 90b53f1

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +41 -26
README.md CHANGED
@@ -1,35 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
1
  # Multimodal AI Doctor – An Agentic AI Project
2
 
3
- **Multimodal AI Doctor** is an **agentic multimodal assistant** built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
4
- It combines **speech, vision, and reasoning** through a series of cooperating LLMs, simulating how a real doctor listens, observes, and responds concisely.
5
- The system integrates **voice input, image analysis, clinical reasoning, and voice output** into a single pipeline.
6
 
7
  ---
8
 
9
  ## Features
10
 
11
- * Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
12
- * Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
13
- * Generate a concise medical-style response (2 sentences maximum, human-like tone)
14
- * Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
15
- * Gradio-based interactive UI
16
 
17
  ---
18
 
19
  ## Project Structure
20
 
21
  ```
 
22
  .
23
  ├── app.py # Gradio UI + main workflow
24
- ├── brain_of_the_doctor.py # Image encoding + Groq multimodal analysis
25
- ├── voice_of_the_patient.py # Audio recording + Groq Whisper transcription
26
- ├── voice_of_the_doctor.py # ElevenLabs + gTTS text-to-speech
27
  ├── requirements.txt # Python dependencies
28
  ├── .env # Environment variables (API keys)
29
- ├── .gitignore # Ignore venv, __pycache__, .env, etc.
30
  ├── images/ # Folder for saving test/sample images
31
  └── README.md # Documentation
32
- ```
 
33
 
34
  ---
35
 
@@ -37,22 +50,22 @@ The system integrates **voice input, image analysis, clinical reasoning, and voi
37
 
38
  The system uses **multiple LLM agents** to process multimodal input step by step:
39
 
40
- 1. **Symptom Agent** – extracts structured meaning from patient speech (via Whisper transcription).
41
- 2. **Vision Agent** – analyzes uploaded medical images (X-ray, MRI, scan).
42
- 3. **Reasoning Agent** – integrates speech and image findings into a medical interpretation.
43
- 4. **Response Agent** – formats the answer in a concise, empathetic, doctor-style tone (≤ 2 sentences).
44
- 5. **Voice Agent** – delivers the response using ElevenLabs (WAV, fallback gTTS).
45
 
46
- This makes the project an **agentic AI pipeline**, where multiple specialized models cooperate to mimic a doctor’s diagnostic process.
47
 
48
  ---
49
 
50
  ## Requirements
51
 
52
- * Python 3.10 or higher
53
- * FFmpeg installed and available in PATH (required by pydub)
54
- * A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
55
- * An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
56
 
57
  ---
58
 
@@ -63,7 +76,7 @@ This makes the project an **agentic AI pipeline**, where multiple specialized mo
63
  ```bash
64
  git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
65
  cd ai-doctor-2.0-voice-and-vision
66
- ```
67
 
68
  2. Create and activate a virtual environment:
69
 
@@ -99,7 +112,7 @@ This makes the project an **agentic AI pipeline**, where multiple specialized mo
99
  Start the Gradio app:
100
 
101
  ```bash
102
- python gradio_app.py
103
  ```
104
 
105
  The app will launch locally at:
@@ -158,4 +171,6 @@ For questions, issues, or collaboration, please contact:
158
  **Email:** [sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)
159
  **LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
160
 
161
- ---
 
 
 
1
+ ---
2
+ title: Multimodal AI Doctor – An Agentic AI Project
3
+ emoji: 🩺
4
+ colorFrom: indigo
5
+ colorTo: blue
6
+ sdk: gradio
7
+ sdk_version: 3.50.2
8
+ app_file: app.py
9
+ pinned: false
10
+ ---
11
+
12
  # Multimodal AI Doctor – An Agentic AI Project
13
 
14
+ **Multimodal AI Doctor** is an **agentic multimodal assistant** built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
15
+ It combines **speech, vision, and reasoning** through a series of cooperating LLMs, simulating how a real doctor listens, observes, and responds concisely.
16
+ The system integrates **voice input, image analysis, clinical reasoning, and voice output** into a single pipeline.
17
 
18
  ---
19
 
20
  ## Features
21
 
22
+ * Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
23
+ * Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
24
+ * Generate a concise medical-style response (2 sentences maximum, human-like tone)
25
+ * Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
26
+ * Gradio-based interactive UI
27
 
28
  ---
29
 
30
  ## Project Structure
31
 
32
  ```
33
+
34
  .
35
  ├── app.py # Gradio UI + main workflow
36
+ ├── brain\_of\_the\_doctor.py # Image encoding + Groq multimodal analysis
37
+ ├── voice\_of\_the\_patient.py # Audio recording + Groq Whisper transcription
38
+ ├── voice\_of\_the\_doctor.py # ElevenLabs + gTTS text-to-speech
39
  ├── requirements.txt # Python dependencies
40
  ├── .env # Environment variables (API keys)
41
+ ├── .gitignore # Ignore venv, **pycache**, .env, etc.
42
  ├── images/ # Folder for saving test/sample images
43
  └── README.md # Documentation
44
+
45
+ ````
46
 
47
  ---
48
 
 
50
 
51
  The system uses **multiple LLM agents** to process multimodal input step by step:
52
 
53
+ 1. **Symptom Agent** – extracts structured meaning from patient speech (via Whisper transcription).
54
+ 2. **Vision Agent** – analyzes uploaded medical images (X-ray, MRI, scan).
55
+ 3. **Reasoning Agent** – integrates speech and image findings into a medical interpretation.
56
+ 4. **Response Agent** – formats the answer in a concise, empathetic, doctor-style tone (≤ 2 sentences).
57
+ 5. **Voice Agent** – delivers the response using ElevenLabs (WAV, fallback gTTS).
58
 
59
+ This makes the project an **agentic AI pipeline**, where multiple specialized models cooperate to mimic a doctor’s diagnostic process.
60
 
61
  ---
62
 
63
  ## Requirements
64
 
65
+ * Python 3.10 or higher
66
+ * FFmpeg installed and available in PATH (required by pydub)
67
+ * A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
68
+ * An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
69
 
70
  ---
71
 
 
76
  ```bash
77
  git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
78
  cd ai-doctor-2.0-voice-and-vision
79
+ ````
80
 
81
  2. Create and activate a virtual environment:
82
 
 
112
  Start the Gradio app:
113
 
114
  ```bash
115
+ python app.py
116
  ```
117
 
118
  The app will launch locally at:
 
171
  **Email:** [sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)
172
  **LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
173
 
174
+ ```
175
+
176
+ ---