Sayeem26s commited on
Commit
90b53f1
·
verified ·
1 Parent(s): 619fb4d

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +160 -145
README.md CHANGED
@@ -1,146 +1,161 @@
1
- # AI Doctor
2
-
3
- **AI Doctor** is a multimodal assistant built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
4
- It allows users to record patient voice, upload medical-related images, and receive a concise **doctor-style spoken response**.
5
-
6
- ---
7
-
8
- ## Features
9
-
10
- * Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
11
- * Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
12
- * Generate a concise medical-style response (2 sentences maximum, human-like tone)
13
- * Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
14
- * Gradio-based interactive UI
15
-
16
- ---
17
-
18
- ## Project Structure
19
-
20
- ```
21
- .
22
- ├── app.py # Gradio UI + main workflow
23
- ├── brain_of_the_doctor.py # Image encoding + Groq multimodal analysis
24
- ├── voice_of_the_patient.py # Audio recording + Groq Whisper transcription
25
- ├── voice_of_the_doctor.py # ElevenLabs + gTTS text-to-speech
26
- ├── requirements.txt # Python dependencies
27
- ├── .env # Environment variables (API keys)
28
- ├── .gitignore # Ignore venv, __pycache__, .env, etc.
29
- ├── images/ # Folder for saving test/sample images
30
- └── README.md # Documentation
31
- ```
32
-
33
- ---
34
-
35
- ## Requirements
36
-
37
- * Python 3.10 or higher
38
- * FFmpeg installed and available in PATH (required by pydub)
39
- * A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
40
- * An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
41
-
42
- ---
43
-
44
- ## Installation
45
-
46
- 1. Clone the repository:
47
-
48
- ```bash
49
- git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
50
- cd ai-doctor-2.0-voice-and-vision
51
- ```
52
-
53
- 2. Create and activate a virtual environment:
54
-
55
- ```bash
56
- python -m venv venv
57
- source venv/bin/activate # Linux/Mac
58
- venv\Scripts\activate # Windows
59
- ```
60
-
61
- 3. Install dependencies:
62
-
63
- ```bash
64
- pip install -r requirements.txt
65
- ```
66
-
67
- 4. Install FFmpeg (if not already installed):
68
-
69
- * Windows: Download from [https://www.gyan.dev/ffmpeg/builds/](https://www.gyan.dev/ffmpeg/builds/) and add `bin/` to PATH
70
- * Linux (Debian/Ubuntu): `sudo apt install ffmpeg`
71
- * macOS (Homebrew): `brew install ffmpeg`
72
-
73
- 5. Create a `.env` file in the project root with your API keys:
74
-
75
- ```
76
- GROQ_API_KEY=your_groq_api_key_here
77
- ELEVEN_API_KEY=your_elevenlabs_api_key_here
78
- ```
79
-
80
- ---
81
-
82
- ## Running the Application
83
-
84
- Start the Gradio app:
85
-
86
- ```bash
87
- python gradio_app.py
88
- ```
89
-
90
- The app will launch locally at:
91
-
92
- ```
93
- http://127.0.0.1:7860
94
- ```
95
-
96
- ---
97
-
98
- ## Usage
99
-
100
- 1. Allow microphone access to record your voice.
101
- 2. Upload a medical image for analysis.
102
- 3. The system will:
103
-
104
- * Transcribe your voice (Whisper Large v3 via Groq)
105
- * Analyze the image + text (Llama 4 Scout via Groq)
106
- * Generate a concise medical-style response
107
- * Play back the response as voice (ElevenLabs or gTTS fallback)
108
-
109
- ---
110
-
111
- ## Models Used
112
-
113
- 1. **Whisper Large v3** (Groq) – Speech-to-Text
114
-
115
- * [Groq API Docs](https://console.groq.com/docs)
116
-
117
- 2. **Llama 4 Scout 17B (Mixture-of-Experts)** (Groq) – Vision-Language reasoning
118
-
119
- * [Groq API Docs](https://console.groq.com/docs)
120
-
121
- 3. **ElevenLabs `eleven_turbo_v2`** – Text-to-Speech (WAV, with MP3 fallback)
122
-
123
- * [ElevenLabs Docs](https://elevenlabs.io/docs)
124
-
125
- 4. **gTTS (Google Text-to-Speech)** – Backup Text-to-Speech
126
-
127
- * [PyPI gTTS](https://pypi.org/project/gTTS/)
128
-
129
- ---
130
-
131
- ## Notes
132
-
133
- * ElevenLabs free-tier accounts may not allow **WAV output** or certain custom voices. In that case, the code automatically falls back to **MP3** output with a safe built-in voice.
134
- * Ensure FFmpeg is correctly installed; otherwise, audio export with pydub will fail.
135
- * Gradio will automatically handle playback of both WAV and MP3 outputs.
136
-
137
- ---
138
-
139
- ## Support
140
-
141
- For questions, issues, or collaboration, please contact:
142
-
143
- **[sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)**
144
- **LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
145
-
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
146
  ---
 
1
+ # Multimodal AI Doctor – An Agentic AI Project
2
+
3
+ **Multimodal AI Doctor** is an **agentic multimodal assistant** built with **Gradio**, **Groq APIs**, and **ElevenLabs**.
4
+ It combines **speech, vision, and reasoning** through a series of cooperating LLMs, simulating how a real doctor listens, observes, and responds concisely.
5
+ The system integrates **voice input, image analysis, clinical reasoning, and voice output** into a single pipeline.
6
+
7
+ ---
8
+
9
+ ## Features
10
+
11
+ * Record patient voice from microphone (Speech-to-Text using **Whisper Large v3** on Groq)
12
+ * Upload an image (diagnosis/medical-related) for analysis (Vision-Language reasoning using **Llama 4 Scout** on Groq)
13
+ * Generate a concise medical-style response (2 sentences maximum, human-like tone)
14
+ * Convert response to voice (Text-to-Speech using **ElevenLabs** with WAV output, fallback to **gTTS** if needed)
15
+ * Gradio-based interactive UI
16
+
17
+ ---
18
+
19
+ ## Project Structure
20
+
21
+ ```
22
+ .
23
+ ├── app.py # Gradio UI + main workflow
24
+ ├── brain_of_the_doctor.py # Image encoding + Groq multimodal analysis
25
+ ├── voice_of_the_patient.py # Audio recording + Groq Whisper transcription
26
+ ├── voice_of_the_doctor.py # ElevenLabs + gTTS text-to-speech
27
+ ├── requirements.txt # Python dependencies
28
+ ├── .env # Environment variables (API keys)
29
+ ├── .gitignore # Ignore venv, __pycache__, .env, etc.
30
+ ├── images/ # Folder for saving test/sample images
31
+ └── README.md # Documentation
32
+ ```
33
+
34
+ ---
35
+
36
+ ## Agentic AI Workflow
37
+
38
+ The system uses **multiple LLM agents** to process multimodal input step by step:
39
+
40
+ 1. **Symptom Agent** – extracts structured meaning from patient speech (via Whisper transcription).
41
+ 2. **Vision Agent** – analyzes uploaded medical images (X-ray, MRI, scan).
42
+ 3. **Reasoning Agent** – integrates speech and image findings into a medical interpretation.
43
+ 4. **Response Agent** – formats the answer in a concise, empathetic, doctor-style tone (≤ 2 sentences).
44
+ 5. **Voice Agent** – delivers the response using ElevenLabs (WAV, fallback gTTS).
45
+
46
+ This makes the project an **agentic AI pipeline**, where multiple specialized models cooperate to mimic a doctor’s diagnostic process.
47
+
48
+ ---
49
+
50
+ ## Requirements
51
+
52
+ * Python 3.10 or higher
53
+ * FFmpeg installed and available in PATH (required by pydub)
54
+ * A Groq API key (obtain from [https://console.groq.com](https://console.groq.com))
55
+ * An ElevenLabs API key (obtain from [https://elevenlabs.io](https://elevenlabs.io))
56
+
57
+ ---
58
+
59
+ ## Installation
60
+
61
+ 1. Clone the repository:
62
+
63
+ ```bash
64
+ git clone https://github.com/your-username/ai-doctor-2.0-voice-and-vision.git
65
+ cd ai-doctor-2.0-voice-and-vision
66
+ ```
67
+
68
+ 2. Create and activate a virtual environment:
69
+
70
+ ```bash
71
+ python -m venv venv
72
+ source venv/bin/activate # Linux/Mac
73
+ venv\Scripts\activate # Windows
74
+ ```
75
+
76
+ 3. Install dependencies:
77
+
78
+ ```bash
79
+ pip install -r requirements.txt
80
+ ```
81
+
82
+ 4. Install FFmpeg (if not already installed):
83
+
84
+ * Windows: [Download builds](https://www.gyan.dev/ffmpeg/builds/) and add `bin/` to PATH
85
+ * Linux (Debian/Ubuntu): `sudo apt install ffmpeg`
86
+ * macOS (Homebrew): `brew install ffmpeg`
87
+
88
+ 5. Create a `.env` file in the project root with your API keys:
89
+
90
+ ```
91
+ GROQ_API_KEY=your_groq_api_key_here
92
+ ELEVEN_API_KEY=your_elevenlabs_api_key_here
93
+ ```
94
+
95
+ ---
96
+
97
+ ## Running the Application
98
+
99
+ Start the Gradio app:
100
+
101
+ ```bash
102
+ python gradio_app.py
103
+ ```
104
+
105
+ The app will launch locally at:
106
+
107
+ ```
108
+ http://127.0.0.1:7860
109
+ ```
110
+
111
+ ---
112
+
113
+ ## Usage
114
+
115
+ 1. Allow microphone access to record your voice.
116
+ 2. Upload a medical image for analysis.
117
+ 3. The system will:
118
+
119
+ * Transcribe your voice (Whisper Large v3 via Groq)
120
+ * Analyze the image + text (Llama 4 Scout via Groq)
121
+ * Generate a concise medical-style response
122
+ * Play back the response as voice (ElevenLabs or gTTS fallback)
123
+
124
+ ---
125
+
126
+ ## Models Used
127
+
128
+ 1. **Whisper Large v3** (Groq) – Speech-to-Text
129
+
130
+ * [Groq API Docs](https://console.groq.com/docs)
131
+
132
+ 2. **Llama 4 Scout 17B (Mixture-of-Experts)** (Groq) – Vision-Language reasoning
133
+
134
+ * [Groq API Docs](https://console.groq.com/docs)
135
+
136
+ 3. **ElevenLabs `eleven_turbo_v2`** – Text-to-Speech (WAV, with MP3 fallback)
137
+
138
+ * [ElevenLabs Docs](https://elevenlabs.io/docs)
139
+
140
+ 4. **gTTS (Google Text-to-Speech)** – Backup Text-to-Speech
141
+
142
+ * [PyPI gTTS](https://pypi.org/project/gTTS/)
143
+
144
+ ---
145
+
146
+ ## Notes
147
+
148
+ * ElevenLabs free-tier accounts may not allow WAV output or certain custom voices. In that case, the code automatically falls back to MP3 output with a safe built-in voice.
149
+ * Ensure FFmpeg is correctly installed; otherwise, audio export with pydub will fail.
150
+ * Gradio will automatically handle playback of both WAV and MP3 outputs.
151
+
152
+ ---
153
+
154
+ ## Support
155
+
156
+ For questions, issues, or collaboration, please contact:
157
+
158
+ **Email:** [sayeem26s@gmail.com](mailto:sayeem26s@gmail.com)
159
+ **LinkedIn:** [https://www.linkedin.com/in/s-m-shahriar-26s/](https://www.linkedin.com/in/s-m-shahriar-26s/)
160
+
161
  ---