DaJulster commited on
Commit
d775234
·
1 Parent(s): f3ebafe

chore: refresh knowledge base

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. docs/faiss/document_lookup.txt +0 -0
  2. docs/faiss/index.faiss +2 -2
  3. docs/faiss/index.pkl +2 -2
  4. docs/faiss/metadata.pkl +2 -2
  5. docs/github_activity/Abudy72_JustIn_PR_1_Added_some_files....md +13 -0
  6. docs/github_activity/GDG-Guelph_Git-Workshop_PR_11_Add_new_user_Julien_Serbanescu_to_data.json.md +13 -0
  7. docs/github_activity/GDG-Guelph_Git-Workshop_PR_12_Line_195_added_a_max_distance_for_push_effect.md +13 -0
  8. docs/github_activity/Julien-ser_Java_PR_1_Mychanges.md +28 -0
  9. docs/github_activity/Julien-ser_agentic-slide-filler_PR_1_wiggum_session.md +58 -0
  10. docs/github_activity/Julien-ser_agentic-slide-filler_PR_2_wiggum_session.md +61 -0
  11. docs/github_activity/Julien-ser_flash-quote-api_PR_1_wiggum_session.md +44 -0
  12. docs/github_activity/Julien-ser_invoice-resolver-ai_PR_1_wiggum_session_1774066153.md +50 -0
  13. docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_1_Marco_branch.md +13 -0
  14. docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_2_Haris_edits_extending_from_punkt.md +13 -0
  15. docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_3_Validation_testing.md +13 -0
  16. docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_1_Add_runme.md_with_picoCTF_flag_from_runme.py.md +13 -0
  17. docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_2_Add_staticnoise.md_with_picoCTF_flag_from_static_binary_analysis.md +13 -0
  18. docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_3_Master.md +13 -0
  19. docs/github_activity/Julien-ser_splunk-mcp-integration_PR_1_Wiggum_session.md +33 -0
  20. docs/github_activity/faizm10_uog-webring_PR_16_Add_user_Julien_Serbanescu.md +50 -0
  21. docs/github_activity/nltk_nltk_PR_3470_Pull_request_for_issue_3413_-_Python_package_for_extra_tokenizers.md +14 -0
  22. docs/github_activity/vllm-project_vllm_PR_35358_Bugfix_Frontend_Fix_reasoning-end_detection_to_check_prompt_tail_o.md +51 -0
  23. docs/papers/arxiv_2510.21983v1_Uncovering_the_Persuasive_Fingerprint_of_LLMs_in_Jailbreaking_Attacks.md +15 -0
  24. docs/pdfs/resume.pdf +0 -3
  25. docs/readmes/AKO-Agentic-Kernel-Optimization-README.md +94 -0
  26. docs/readmes/Free-Wiggum-README.md +310 -0
  27. docs/readmes/Git-Workshop-README.md +23 -0
  28. docs/readmes/Our-Papers-README.md +15 -0
  29. docs/readmes/ResumeWorthy-README.md +1 -1
  30. docs/readmes/Stockmetricagent-README.md +154 -0
  31. docs/readmes/WiggumLoopAgenticSWDeveloper-README.md +526 -0
  32. docs/readmes/agentic-founders-finding-README.md +201 -0
  33. docs/readmes/agentic-slide-filler-README.md +148 -0
  34. docs/readmes/agentic-team-README.md +139 -0
  35. docs/readmes/basestation_server-README.md +1 -1
  36. docs/readmes/calorie-counter-README.md +116 -0
  37. docs/readmes/causal-github-pr-analysis-README.md +38 -0
  38. docs/readmes/causal-model-README.md +137 -0
  39. docs/readmes/cuda-optimizer-README.md +231 -0
  40. docs/readmes/databricks-powerbi-pipeline-README.md +209 -0
  41. docs/readmes/edgebot-ai-README.md +815 -0
  42. docs/readmes/esp-robot-README.md +281 -0
  43. docs/readmes/flash-quote-api-README.md +142 -0
  44. docs/readmes/football-management-sim-README.md +577 -0
  45. docs/readmes/hipstercheck-README.md +531 -0
  46. docs/readmes/internlearningnetwork-README.md +327 -0
  47. docs/readmes/invoice-resolver-ai-README.md +427 -0
  48. docs/readmes/jira-and-confluence-agent-README.md +306 -0
  49. docs/readmes/julien-rag-README.md +298 -0
  50. docs/readmes/mlflow-ai-experiment-README.md +379 -0
docs/faiss/document_lookup.txt CHANGED
The diff for this file is too large to render. See raw diff
 
docs/faiss/index.faiss CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4345d04c2c185f8c9c3cbd3106b7d7c379f08df78b1d9dfd0cf5cd66fe03fc3f
3
- size 1048621
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:761c3e1c863aed9b8194abb693fb5cd07c237ebf4987efd4f58ab4349df6cbd7
3
+ size 3424301
docs/faiss/index.pkl CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ee3a2de1ce31b6c7d0162e15e6cac52ba5ad314d3062c93b89d360014a547cbf
3
- size 239138
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8a2bb4a67ed657a46ea93c309f4f686b6c7115bd2b18c7605c5b28c13d686ff1
3
+ size 767874
docs/faiss/metadata.pkl CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e21a8afdd183583c8da91ea779e9e50deeff9e4cb94fa0a221743e890cdc7ac5
3
- size 4461
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:16da6051eb9ab58343238c7c7dd50c32229abdfab355d06871f7d817904ef93d
3
+ size 16460
docs/github_activity/Abudy72_JustIn_PR_1_Added_some_files....md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Added some files...
2
+
3
+ **Repository:** Abudy72/JustIn
4
+ **Number:** #1
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2024-05-04T20:06:00Z
8
+ **Updated:** 2024-05-04T20:06:11Z
9
+ **URL:** https://github.com/Abudy72/JustIn/pull/1
10
+
11
+ ---
12
+
13
+
docs/github_activity/GDG-Guelph_Git-Workshop_PR_11_Add_new_user_Julien_Serbanescu_to_data.json.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Add new user Julien Serbanescu to data.json
2
+
3
+ **Repository:** GDG-Guelph/Git-Workshop
4
+ **Number:** #11
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2026-02-21T01:26:27Z
8
+ **Updated:** 2026-02-21T01:56:20Z
9
+ **URL:** https://github.com/GDG-Guelph/Git-Workshop/pull/11
10
+
11
+ ---
12
+
13
+
docs/github_activity/GDG-Guelph_Git-Workshop_PR_12_Line_195_added_a_max_distance_for_push_effect.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Line 195, added a max distance for push effect
2
+
3
+ **Repository:** GDG-Guelph/Git-Workshop
4
+ **Number:** #12
5
+ **State:** open
6
+ **Labels:** None
7
+ **Created:** 2026-02-21T04:51:19Z
8
+ **Updated:** 2026-02-21T04:51:19Z
9
+ **URL:** https://github.com/GDG-Guelph/Git-Workshop/pull/12
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_Java_PR_1_Mychanges.md ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Mychanges
2
+
3
+ **Repository:** Julien-ser/Java
4
+ **Number:** #1
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-10-20T13:41:55Z
8
+ **Updated:** 2025-10-20T13:43:31Z
9
+ **URL:** https://github.com/Julien-ser/Java/pull/1
10
+
11
+ ---
12
+
13
+ <!--
14
+ Thank you for your contribution!
15
+ In order to reduce the number of notifications sent to the maintainers, please:
16
+ - create your PR as draft, cf. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/about-pull-requests#draft-pull-requests,
17
+ - make sure that all of the CI checks pass,
18
+ - mark your PR as ready for review, cf. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/changing-the-stage-of-a-pull-request#marking-a-pull-request-as-ready-for-review
19
+ -->
20
+
21
+ <!-- For completed items, change [ ] to [x] -->
22
+
23
+ - [ ] I have read [CONTRIBUTING.md](https://github.com/TheAlgorithms/Java/blob/master/CONTRIBUTING.md).
24
+ - [ ] This pull request is all my own work -- I have not plagiarized it.
25
+ - [ ] All filenames are in PascalCase.
26
+ - [ ] All functions and variable names follow Java naming conventions.
27
+ - [ ] All new algorithms have a URL in their comments that points to Wikipedia or other similar explanations.
28
+ - [ ] All new code is formatted with `clang-format -i --style=file path/to/your/file.java`
docs/github_activity/Julien-ser_agentic-slide-filler_PR_1_wiggum_session.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: wiggum/session
2
+
3
+ **Repository:** Julien-ser/agentic-slide-filler
4
+ **Number:** #1
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2026-03-29T02:51:40Z
8
+ **Updated:** 2026-03-29T02:59:41Z
9
+ **URL:** https://github.com/Julien-ser/agentic-slide-filler/pull/1
10
+
11
+ ---
12
+
13
+ - **docs: define technical architecture and system design**
14
+ - **Iteration 1: Define technical architecture: select Python as base language, choose `python-pptx` for PPT manipulation, `openai` API (or `anthropic`) for AI generation, and `python-docx` for Word outline parsing; create system architecture diagram**
15
+ - **feat: setup project structure and dependencies**
16
+ - **Iteration 2: Set up project structure with `src/`, `tests/`, `templates/`, `output/`, `docs/` directories; initialize Python virtual environment and `requirements.txt` with `python-pptx>=0.6.21`, `openai>=1.0.0`, `python-docx>=0.8.11`**
17
+ - **Iteration 3: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
18
+ - **Iteration 4: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
19
+ - **chore: mark configuration management task as complete**
20
+ - **Iteration 5: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
21
+ - **docs: mark Git setup and README as complete**
22
+ - **Iteration 6: Establish Git repository with `.gitignore` for Python, IDE files, and sensitive configs; write initial `README.md` with project overview and setup instructions**
23
+ - **Iteration 7: Build PPT template parser: create `TemplateParser` class in `src/template_parser.py` that reads `.pptx` files, extracts slide layouts, placeholders, and shape types, returns structured template metadata**
24
+ - **test: add comprehensive unit tests for TemplateParser**
25
+ - **fix: add src/__init__.py for package imports**
26
+ - **Iteration 8: Build PPT template parser: create `TemplateParser` class in `src/template_parser.py` that reads `.pptx` files, extracts slide layouts, placeholders, and shape types, returns structured template metadata**
27
+ - **Iteration 9: Implement doc outline parser: develop `OutlineParser` class in `src/outline_parser.py` supporting `.docx` and `.md` formats; extracts headings, sections, and content hierarchy into JSON structure**
28
+ - **feat: add ContentMapper class for section-to-placeholder alignment**
29
+ - **Iteration 10: Create content mapping engine: build `ContentMapper` class in `src/content_mapper.py` that aligns outline sections with template slide placeholders using section titles and placeholder tags**
30
+ - **Iteration 12: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
31
+ - **Fix template_parser: dynamic PP_PLACEHOLDER mapping to handle missing enum attributes, and robust test fixture for placeholder indices**
32
+ - **Iteration 13: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
33
+ - **fix: handle KeyError when extracting placeholder indices in TemplateParser.\n\nSome slide layouts contain placeholders that reference invalid indices,\ncausing a KeyError from python-pptx when accessing placeholder_format.idx.\nThe _extract_placeholders method now catches KeyError and skips such\nplaceholders gracefully.**
34
+ - **Iteration 14: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
35
+ - **fix: resolve lint and type errors in template_parser.py**
36
+ - **feat(ai): complete Phase 3 AI tasks; add prompts, generator, validator, and fallback**
37
+ - **Iteration 15: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
38
+ - **Iteration 16: Assemble end-to-end pipeline: create main `SlideFiller` class in `src/slide_filler.py` that orchestrates parsing, mapping, generation, and filling; integrate all modules with dependency injection**
39
+
40
+
41
+ <!-- This is an auto-generated comment: release notes by coderabbit.ai -->
42
+
43
+ ## Summary by CodeRabbit
44
+
45
+ ## Release Notes
46
+
47
+ * **New Features**
48
+ * Complete command-line interface for converting document outlines into PowerPoint presentations.
49
+ * Support for multiple document formats (DOCX, Markdown) and mapping strategies.
50
+ * LLM-powered content generation with multiple model support and automatic fallback.
51
+ * Content validation and caching to optimize generation performance.
52
+ * Dry-run mode for testing workflows without producing output.
53
+
54
+ * **Documentation**
55
+ * Comprehensive setup and usage guide with system architecture overview.
56
+ * Configuration examples and environment variable templates provided.
57
+
58
+ <!-- end of auto-generated comment: release notes by coderabbit.ai -->
docs/github_activity/Julien-ser_agentic-slide-filler_PR_2_wiggum_session.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: wiggum/session
2
+
3
+ **Repository:** Julien-ser/agentic-slide-filler
4
+ **Number:** #2
5
+ **State:** open
6
+ **Labels:** None
7
+ **Created:** 2026-03-29T02:56:22Z
8
+ **Updated:** 2026-03-29T03:05:45Z
9
+ **URL:** https://github.com/Julien-ser/agentic-slide-filler/pull/2
10
+
11
+ ---
12
+
13
+ - **docs: define technical architecture and system design**
14
+ - **Iteration 1: Define technical architecture: select Python as base language, choose `python-pptx` for PPT manipulation, `openai` API (or `anthropic`) for AI generation, and `python-docx` for Word outline parsing; create system architecture diagram**
15
+ - **feat: setup project structure and dependencies**
16
+ - **Iteration 2: Set up project structure with `src/`, `tests/`, `templates/`, `output/`, `docs/` directories; initialize Python virtual environment and `requirements.txt` with `python-pptx>=0.6.21`, `openai>=1.0.0`, `python-docx>=0.8.11`**
17
+ - **Iteration 3: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
18
+ - **Iteration 4: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
19
+ - **chore: mark configuration management task as complete**
20
+ - **Iteration 5: Create configuration management: implement `config.yaml` for API keys, template paths, and output settings; add `.env.example` template with `OPENAI_API_KEY` placeholder**
21
+ - **docs: mark Git setup and README as complete**
22
+ - **Iteration 6: Establish Git repository with `.gitignore` for Python, IDE files, and sensitive configs; write initial `README.md` with project overview and setup instructions**
23
+ - **Iteration 7: Build PPT template parser: create `TemplateParser` class in `src/template_parser.py` that reads `.pptx` files, extracts slide layouts, placeholders, and shape types, returns structured template metadata**
24
+ - **test: add comprehensive unit tests for TemplateParser**
25
+ - **fix: add src/__init__.py for package imports**
26
+ - **Iteration 8: Build PPT template parser: create `TemplateParser` class in `src/template_parser.py` that reads `.pptx` files, extracts slide layouts, placeholders, and shape types, returns structured template metadata**
27
+ - **Iteration 9: Implement doc outline parser: develop `OutlineParser` class in `src/outline_parser.py` supporting `.docx` and `.md` formats; extracts headings, sections, and content hierarchy into JSON structure**
28
+ - **feat: add ContentMapper class for section-to-placeholder alignment**
29
+ - **Iteration 10: Create content mapping engine: build `ContentMapper` class in `src/content_mapper.py` that aligns outline sections with template slide placeholders using section titles and placeholder tags**
30
+ - **Iteration 12: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
31
+ - **Fix template_parser: dynamic PP_PLACEHOLDER mapping to handle missing enum attributes, and robust test fixture for placeholder indices**
32
+ - **Iteration 13: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
33
+ - **fix: handle KeyError when extracting placeholder indices in TemplateParser.\n\nSome slide layouts contain placeholders that reference invalid indices,\ncausing a KeyError from python-pptx when accessing placeholder_format.idx.\nThe _extract_placeholders method now catches KeyError and skips such\nplaceholders gracefully.**
34
+ - **Iteration 14: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
35
+ - **fix: resolve lint and type errors in template_parser.py**
36
+ - **feat(ai): complete Phase 3 AI tasks; add prompts, generator, validator, and fallback**
37
+ - **Iteration 15: Add template validation: write unit tests in `tests/test_template_parser.py` that verify placeholder detection, layout identification, and error handling for corrupted templates**
38
+ - **Iteration 16: Assemble end-to-end pipeline: create main `SlideFiller` class in `src/slide_filler.py` that orchestrates parsing, mapping, generation, and filling; integrate all modules with dependency injection**
39
+ - **Trigger CodeRabbit review**
40
+ - **Trigger CodeRabbit review**
41
+ - **Trigger CodeRabbit review**
42
+ - **Trigger CodeRabbit review**
43
+
44
+
45
+ <!-- This is an auto-generated comment: release notes by coderabbit.ai -->
46
+
47
+ ## Summary by CodeRabbit
48
+
49
+ * **New Features**
50
+ * Added a complete end-to-end tool for generating PowerPoint presentations from document outlines using LLMs.
51
+ * Added command-line interface with support for templates, outlines (`.docx`, `.md`), configuration options, and execution modes.
52
+ * Added integration with OpenAI and Anthropic APIs for AI-powered content generation.
53
+ * Added caching and retry mechanisms for efficient API usage.
54
+
55
+ * **Documentation**
56
+ * Updated README with technical architecture, data flow, setup instructions, and configuration details.
57
+
58
+ * **Tests**
59
+ * Added comprehensive test suites for all core components.
60
+
61
+ <!-- end of auto-generated comment: release notes by coderabbit.ai -->
docs/github_activity/Julien-ser_flash-quote-api_PR_1_wiggum_session.md ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: wiggum/session
2
+
3
+ **Repository:** Julien-ser/flash-quote-api
4
+ **Number:** #1
5
+ **State:** open
6
+ **Labels:** None
7
+ **Created:** 2026-03-27T22:54:24Z
8
+ **Updated:** 2026-03-27T23:05:01Z
9
+ **URL:** https://github.com/Julien-ser/flash-quote-api/pull/1
10
+
11
+ ---
12
+
13
+ - **Iteration 1: Define the quote data structure (e.g., JSON array with fields: id, text, author, category)**
14
+ - **feat: define quote data structure with JSON format**
15
+ - **Iteration 2: Define the quote data structure (e.g., JSON array with fields: id, text, author, category)**
16
+ - **Iteration 3: Choose storage method: in-memory array vs JSON file vs SQLite database**
17
+ - **chore: initialize project structure and mark Phase 1 complete**
18
+ - **Iteration 5: Create quotes_data.json with at least 50 inspirational quotes**
19
+ - **Iteration 7: Generate OpenAPI/Swagger docs with FastAPI automatic docs**
20
+ - **Iteration 8: Create Dockerfile for containerized deployment**
21
+ - **feat: add Dockerfile for containerized deployment**
22
+ - **Iteration 9: Create Dockerfile for containerized deployment**
23
+ - **ci: fix workflow to use requirements.txt and mark task complete**
24
+ - **Iteration 10: Set up GitHub Actions for CI: linting (ruff), testing on push**
25
+ - **Final worker session push**
26
+ - **Finalize PR: all tasks complete**
27
+
28
+
29
+ <!-- This is an auto-generated comment: release notes by coderabbit.ai -->
30
+
31
+ ## Summary by CodeRabbit
32
+
33
+ * **New Features**
34
+ * Introduced flash-quote-api service with three endpoints: retrieve random quotes, list all quotes with optional category filtering, and fetch quotes by ID.
35
+ * Enabled CORS for cross-origin requests.
36
+
37
+ * **Documentation**
38
+ * Added interactive API documentation via Swagger/OpenAPI.
39
+
40
+ * **Chores**
41
+ * Added Docker support for containerized deployment.
42
+ * Configured GitHub Actions for automated testing and CI.
43
+
44
+ <!-- end of auto-generated comment: release notes by coderabbit.ai -->
docs/github_activity/Julien-ser_invoice-resolver-ai_PR_1_wiggum_session_1774066153.md ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: wiggum/session 1774066153
2
+
3
+ **Repository:** Julien-ser/invoice-resolver-ai
4
+ **Number:** #1
5
+ **State:** open
6
+ **Labels:** None
7
+ **Created:** 2026-03-21T06:55:42Z
8
+ **Updated:** 2026-03-22T03:46:39Z
9
+ **URL:** https://github.com/Julien-ser/invoice-resolver-ai/pull/1
10
+
11
+ ---
12
+
13
+ - **feat: complete Task 1.2 - initialize FastAPI project structure**
14
+ - **docs: add comprehensive API documentation (Task 1.4)**
15
+ - **docs: mark Task 2.2 complete; update Python version note; bump pydantic to 2.11.0**
16
+ - **feat: implement email automation with Jinja2 templates and Celery tasks**
17
+ - **feat: implement AI dispute letter drafting with OpenAI and Anthropic**
18
+ - **feat: add PDF generation with WeasyPrint for dispute letters and small claims forms**
19
+ - **feat: add Streamlit dashboard with JWT authentication and user interface**
20
+ - **feat: add admin panel to Streamlit dashboard with user management and system health monitoring**
21
+ - **feat: implement Stripe Billing integration with subscription management**
22
+ - **docs: add Billing Integration section to README**
23
+ - **chore: mark Task 5.1 as complete - comprehensive pytest test suite in place\n\nTests cover email, AI, PDF, integrations, auth, invoices, webhooks, billing, A/B testing, admin, config, middleware, and Celery tasks. Requires PostgreSQL and Redis to run (see docker-compose.yml).**
24
+ - **ci: add Docker multi-stage builds and production deployment configuration**
25
+ - **Final worker session push**
26
+
27
+
28
+ <!-- This is an auto-generated comment: release notes by coderabbit.ai -->
29
+
30
+ ## Summary by CodeRabbit
31
+
32
+ * **New Features**
33
+ * Added invoice management REST API with JWT authentication
34
+ * Integrated payment provider support (Stripe, PayPal, Plaid) with webhook event handling
35
+ * Added automated email follow-up sequences for invoice reminders
36
+ * Added admin dashboard for system monitoring and user management
37
+ * Added containerized deployment infrastructure with Docker and Fly.io support
38
+
39
+ * **Documentation**
40
+ * Added comprehensive API reference documentation
41
+ * Added deployment runbook and database schema documentation
42
+
43
+ * **Tests**
44
+ * Added test coverage for authentication, invoices, and webhooks
45
+
46
+ * **Chores**
47
+ * Added monitoring stack with Prometheus and Grafana integration
48
+ * Added CI/CD pipeline configuration
49
+
50
+ <!-- end of auto-generated comment: release notes by coderabbit.ai -->
docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_1_Marco_branch.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Marco branch
2
+
3
+ **Repository:** Julien-ser/nltk_Group1_ENGG4450
4
+ **Number:** #1
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-11-22T17:44:33Z
8
+ **Updated:** 2025-11-22T17:44:50Z
9
+ **URL:** https://github.com/Julien-ser/nltk_Group1_ENGG4450/pull/1
10
+
11
+ ---
12
+
13
+ Marco's edits
docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_2_Haris_edits_extending_from_punkt.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Haris' edits, extending from punkt
2
+
3
+ **Repository:** Julien-ser/nltk_Group1_ENGG4450
4
+ **Number:** #2
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-11-22T17:47:14Z
8
+ **Updated:** 2025-11-22T17:47:40Z
9
+ **URL:** https://github.com/Julien-ser/nltk_Group1_ENGG4450/pull/2
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_nltk_Group1_ENGG4450_PR_3_Validation_testing.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Validation testing
2
+
3
+ **Repository:** Julien-ser/nltk_Group1_ENGG4450
4
+ **Number:** #3
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-11-25T02:34:39Z
8
+ **Updated:** 2025-11-25T02:34:50Z
9
+ **URL:** https://github.com/Julien-ser/nltk_Group1_ENGG4450/pull/3
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_1_Add_runme.md_with_picoCTF_flag_from_runme.py.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Add runme.md with picoCTF flag from runme.py
2
+
3
+ **Repository:** Julien-ser/picoCTFopencodewriteups
4
+ **Number:** #1
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-10-29T17:46:01Z
8
+ **Updated:** 2025-10-29T17:46:18Z
9
+ **URL:** https://github.com/Julien-ser/picoCTFopencodewriteups/pull/1
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_2_Add_staticnoise.md_with_picoCTF_flag_from_static_binary_analysis.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Add staticnoise.md with picoCTF flag from static binary analysis
2
+
3
+ **Repository:** Julien-ser/picoCTFopencodewriteups
4
+ **Number:** #2
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-10-29T17:50:16Z
8
+ **Updated:** 2025-10-29T17:50:26Z
9
+ **URL:** https://github.com/Julien-ser/picoCTFopencodewriteups/pull/2
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_picoCTFopencodewriteups_PR_3_Master.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Master
2
+
3
+ **Repository:** Julien-ser/picoCTFopencodewriteups
4
+ **Number:** #3
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2025-10-30T00:43:02Z
8
+ **Updated:** 2025-10-30T00:43:14Z
9
+ **URL:** https://github.com/Julien-ser/picoCTFopencodewriteups/pull/3
10
+
11
+ ---
12
+
13
+
docs/github_activity/Julien-ser_splunk-mcp-integration_PR_1_Wiggum_session.md ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Wiggum/session
2
+
3
+ **Repository:** Julien-ser/splunk-mcp-integration
4
+ **Number:** #1
5
+ **State:** open
6
+ **Labels:** None
7
+ **Created:** 2026-03-29T02:40:02Z
8
+ **Updated:** 2026-03-29T02:45:20Z
9
+ **URL:** https://github.com/Julien-ser/splunk-mcp-integration/pull/1
10
+
11
+ ---
12
+
13
+
14
+
15
+ <!-- This is an auto-generated comment: release notes by coderabbit.ai -->
16
+
17
+ ## Summary by CodeRabbit
18
+
19
+ ## Release Notes
20
+
21
+ * **New Features**
22
+ * Introduced Splunk MCP integration with server, client, and tool wrappers for search, indexing, system operations, and AI-powered analysis (SAIA).
23
+ * Added agent orchestration layer with message bus and communication protocol.
24
+
25
+ * **Documentation**
26
+ * Comprehensive project README with setup instructions, architecture overview, and tool descriptions.
27
+ * Project structure and task management documentation.
28
+
29
+ * **Chores**
30
+ * Added development environment configuration (Makefile, Python version specification, environment templates).
31
+ * Updated project gitignore patterns and setup files.
32
+
33
+ <!-- end of auto-generated comment: release notes by coderabbit.ai -->
docs/github_activity/faizm10_uog-webring_PR_16_Add_user_Julien_Serbanescu.md ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Add user Julien Serbanescu
2
+
3
+ **Repository:** faizm10/uog-webring
4
+ **Number:** #16
5
+ **State:** closed
6
+ **Labels:** None
7
+ **Created:** 2026-02-24T15:49:05Z
8
+ **Updated:** 2026-02-24T17:18:05Z
9
+ **URL:** https://github.com/faizm10/uog-webring/pull/16
10
+
11
+ ---
12
+
13
+ ## Member Submission
14
+
15
+ Please fill this out so we can quickly review and merge your addition.
16
+
17
+ ### Member Data
18
+
19
+ - Name: Julien Serbanescu
20
+ - Year: 3rd
21
+ - Website (full URL, https://...): https://julien-ser.github.io/JulienSerbanescu/
22
+ - Role (optional, e.g. job title, "Student", "Founder"): Student Cloud Engineer
23
+
24
+ Rest is on there
25
+
26
+ ### JSON Snippet Added to `data/members.json`
27
+
28
+ Add your entry to the `sites` array in `data/members.json`. Use empty string `""` for any optional link you don't have. Omit `role` if you prefer not to include it.
29
+
30
+ ```json
31
+ {
32
+ "name": "Julien Serbanescu",
33
+ "website": "https://julien-ser.github.io/JulienSerbanescu/",
34
+ "year": 2028,
35
+ "role": "Cloud Engineer @ Co-operators and AI Researcher",
36
+ "links": {
37
+ "twitter": "https://x.com/Da_Julster",
38
+ "linkedin": "https://www.linkedin.com/in/julien-serbanescu-6ba52a241/",
39
+ "github": "https://github.com/Julien-ser"
40
+ }
41
+ ```
42
+
43
+ ### Checklist
44
+
45
+ - [x] I added my entry to `data/members.json` in the `sites` array.
46
+ - [X] My website URL is valid and starts with `https://`.
47
+ - [X] If included, social links are full URLs; unused links use `""`.
48
+ - [X] I kept the JSON valid (no trailing commas, correct quotes).
49
+ - [X] I verified my row appears correctly in the table.
50
+
docs/github_activity/nltk_nltk_PR_3470_Pull_request_for_issue_3413_-_Python_package_for_extra_tokenizers.md ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: Pull request for issue #3413 -> Python package for extra tokenizers
2
+
3
+ **Repository:** nltk/nltk
4
+ **Number:** #3470
5
+ **State:** closed
6
+ **Labels:** tokenizer
7
+ **Created:** 2025-11-25T03:07:44Z
8
+ **Updated:** 2025-11-25T04:49:04Z
9
+ **URL:** https://github.com/nltk/nltk/pull/3470
10
+
11
+ ---
12
+
13
+ Created a new python package to install additional tokenizers instead of the nltk.downloader() class.
14
+ More information and documentation can be seen in the pypi page for this library https://pypi.org/project/nltk-extratokenizers/0.1.2/
docs/github_activity/vllm-project_vllm_PR_35358_Bugfix_Frontend_Fix_reasoning-end_detection_to_check_prompt_tail_o.md ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PR: [Bugfix][Frontend] Fix reasoning-end detection to check prompt tail o…
2
+
3
+ **Repository:** vllm-project/vllm
4
+ **Number:** #35358
5
+ **State:** open
6
+ **Labels:** bug, frontend
7
+ **Created:** 2026-02-26T03:58:27Z
8
+ **Updated:** 2026-02-27T16:25:46Z
9
+ **URL:** https://github.com/vllm-project/vllm/pull/35358
10
+
11
+ ---
12
+
13
+ [Bugfix][Frontend] Fix reasoning-end detection to check prompt tail only (Fix #35349)
14
+
15
+ ## Purpose
16
+
17
+ Fixes #35349.
18
+
19
+ This PR fixes reasoning-end detection in OpenAI chat completion serving so we only treat the prompt as ended reasoning when the **last prompt token** is the reasoning end token (instead of matching anywhere in the prompt).
20
+ It also adds a small helper to safely handle empty/`None` prompt token IDs.
21
+
22
+ ## Test Plan
23
+
24
+ - Run focused unit tests for the new behavior:
25
+ - `python -m pytest tests/entrypoints/openai/test_serving_chat.py -k PromptEndsReasoning -q`
26
+ - (Optional) Run the full serving chat tests:
27
+ - `python -m pytest tests/entrypoints/openai/test_serving_chat.py -q`
28
+
29
+ ## Test Result
30
+
31
+ - Added targeted unit tests covering:
32
+ - end token in the middle of prompt → False
33
+ - end token at prompt tail → True
34
+ - empty / None prompt token IDs → False
35
+ - Local run status: not executed in this environment yet.
36
+ - To finalize before merge, run and paste output from:
37
+ - `python -m pytest tests/entrypoints/openai/test_serving_chat.py -k PromptEndsReasoning -q`
38
+
39
+ ---
40
+ <details>
41
+ <summary> Essential Elements of an Effective PR Description Checklist </summary>
42
+
43
+ - [x] The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
44
+ - [x] The test plan, such as providing test command.
45
+ - [x] The test results, such as pasting the results comparison before and after, or e2e results
46
+ - [x] (Optional) The necessary documentation update was considered (not needed for this internal bug fix).
47
+ - [x] (Optional) Release notes update was considered (not required unless this behavior change is considered user-facing).
48
+
49
+ </details>
50
+
51
+
docs/papers/arxiv_2510.21983v1_Uncovering_the_Persuasive_Fingerprint_of_LLMs_in_Jailbreaking_Attacks.md ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
2
+
3
+ **Authors:** Havva Alizadeh Noughabi, Julien Serbanescu, Fattane Zarrinkalam, Ali Dehghantanha
4
+ **arXiv ID:** 2510.21983v1
5
+ **Published:** 2025-10-24
6
+ **Updated:** 2025-10-24
7
+ **Categories:** cs.CL, cs.AI
8
+ **PDF:** https://arxiv.org/pdf/2510.21983v1
9
+ **Abstract URL:** https://arxiv.org/abs/2510.21983v1
10
+
11
+ ---
12
+
13
+ ## Abstract
14
+
15
+ Despite recent advances, Large Language Models remain vulnerable to jailbreak attacks that bypass alignment safeguards and elicit harmful outputs. While prior research has proposed various attack strategies differing in human readability and transferability, little attention has been paid to the linguistic and psychological mechanisms that may influence a model's susceptibility to such attacks. In this paper, we examine an interdisciplinary line of research that leverages foundational theories of persuasion from the social sciences to craft adversarial prompts capable of circumventing alignment constraints in LLMs. Drawing on well-established persuasive strategies, we hypothesize that LLMs, having been trained on large-scale human-generated text, may respond more compliantly to prompts with persuasive structures. Furthermore, we investigate whether LLMs themselves exhibit distinct persuasive fingerprints that emerge in their jailbreak responses. Empirical evaluations across multiple aligned LLMs reveal that persuasion-aware prompts significantly bypass safeguards, demonstrating their potential to induce jailbreak behaviors. This work underscores the importance of cross-disciplinary insight in addressing the evolving challenges of LLM safety. The code and data are available.
docs/pdfs/resume.pdf DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:540c654c8072993a853f7a059fed4c50debfb4086b4b45ce87ecd5c0e4ca35ea
3
- size 170230
 
 
 
 
docs/readmes/AKO-Agentic-Kernel-Optimization-README.md ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # AKO-Agentic-Kernel-Optimization
2
+ **Repository:** Julien-ser/AKO-Agentic-Kernel-Optimization
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** A proposed system that autonomously generates, profiles and refines GPU kernels using Agentic AI and LLMs in order to optimize high performance AI inference on AMD hardware.
6
+ **Language:** Python
7
+ **Last Updated:** 2026-02-07T01:42:48Z
8
+
9
+ ---
10
+
11
+
12
+ # AKO: Agentic Kernel Optimization for High-Performance Inference on AMD Instinct
13
+
14
+ ## Overview
15
+
16
+
17
+ AKO (Agentic Kernel Optimization) is an autonomous system for optimizing GPU kernel performance on AMD Instinct hardware. By leveraging Agentic AI and Large Language Models (LLMs), AKO automates the generation, profiling, and refinement of GPU kernels—specifically using AMD's TileLang domain-specific language—targeting the CDNA-3/4 architecture (MI300/350 series).
18
+
19
+ The core innovation is a closed "Compile-Profile-Refine" agentic loop, moving beyond manual tuning or brute-force grid searches. The system is designed to unlock the full potential of AMD Instinct GPUs for high-performance AI inference workloads.
20
+
21
+ ## System Workflow
22
+
23
+ 1. **High-Level Specification:** The process begins with a high-level kernel specification (e.g., matrix multiplication or other AI-relevant GPU tasks).
24
+ 2. **Agentic LLM Code Generation:** An LLM-based agent generates specialized GPU kernel code in TileLang, tailored to the given task and hardware.
25
+ 3. **Compilation & Deployment:** The ROCm toolchain (HIP/LLVM) compiles the generated code and deploys it to the MI300X GPU.
26
+ 4. **Profiling & Metrics Collection:** During execution, ROCm's `rocprofiler` collects detailed metrics: execution time, memory bandwidth, utilization, occupancy, cache hit rates, and hardware stalls.
27
+ 5. **Performance Analysis & Feedback:** These metrics are compared to the AITER baseline. Bottlenecks and inefficiencies are identified and reported back to the LLM agent.
28
+ 6. **Reinforcement Learning Loop:** Using an RL system (e.g., verl), the agent receives reward signals based on performance improvements, conditioning it to generate increasingly optimized kernel code.
29
+
30
+ This loop continues iteratively, enabling the agent to autonomously explore the optimization space (tiling sizes, memory layouts, MFMA scheduling, etc.) and converge on high-performance solutions.
31
+
32
+ ## Problem Space
33
+
34
+ Current inference libraries (e.g., vLLM, SGLang) rely on hand-tuned or basic autotuned kernels. TileLang has demonstrated up to 5x speedups over Triton on AMD hardware, but the optimization space is vast and complex. AKO's agentic approach enables:
35
+
36
+ - Automated navigation of kernel design choices (tiling, LDS usage, MFMA scheduling)
37
+ - Hardware-aware code generation and profiling
38
+ - Continuous improvement via RL-driven feedback
39
+
40
+ ## Research Goals
41
+
42
+ - **Autonomous TileLang Synthesis:** Develop an LLM agent that generates TileLang code from high-level mathematical specs.
43
+ - **Closed-Loop Profiling Feedback:** Integrate ROCm's `rocprofiler` to provide performance metrics as reward signals for iterative improvement.
44
+ - **Inference Bottleneck Analysis:** Use Causal AI to model and optimize latency trade-offs in vLLM's PagedAttention vs. SGLang's RadixAttention on AMD Infinity Fabric.
45
+
46
+
47
+ ## Technical Stack
48
+
49
+ - **Languages:** Python (integration), C++/HIP (kernel level), TileLang (DSL)
50
+ - **Frameworks:** `verl` (RL agent training), `AITER` (AMD AI Inference Toolkit)
51
+ - **Hardware:** Instinct MI300X, MI325X, MI350/355 series
52
+
53
+ ## Project Structure
54
+
55
+ ```
56
+ AKO_Project/
57
+ ├── src/
58
+ │ ├── agentic_loop.py # Orchestrates the main optimization loop
59
+ │ ├── kernel_generator.py # LLM agent for kernel code generation
60
+ │ ├── profiling_feedback.py # Performance metric analysis from rocprof
61
+ │ └── hardware_interface.py # AMD hardware interaction (compilation, execution, profiling)
62
+ ├── diagrams/
63
+ │ └── ako_architecture.mmd # Mermaid diagram of system architecture
64
+ └── README.md # This document
65
+ ```
66
+
67
+ ## System Architecture Diagram
68
+
69
+
70
+ Below is the core system architecture, visualized in Mermaid:
71
+
72
+ ```mermaid
73
+ graph TD
74
+ A[High-Level Kernel Specification] --> B(LLM Agent - KernelGenerator)
75
+ subgraph Agentic Loop
76
+ B -- Generates TileLang Code --> C{AMD ROCm Toolchain}
77
+ C -- Compiles & Deploys --> D[Executable Kernel on MI300X]
78
+ D -- Executes & Profiles --> E(ROCm Profiler - rocprof)
79
+ E -- Raw Metrics --> F(Performance Analysis & AITER Comparison)
80
+ F -- Optimization Feedback & Reward Signal --> B
81
+ end
82
+ G[AITER Baseline Library] -. Gold Standard Benchmarks .-> F
83
+ style A fill:#fde7f3,stroke:#333,stroke-width:2px,color:#111
84
+ style B fill:#e6ecff,stroke:#333,stroke-width:2px,color:#111
85
+ style C fill:#eef2ff,stroke:#333,stroke-width:2px,color:#111
86
+ style D fill:#eaf7ea,stroke:#333,stroke-width:2px,color:#111
87
+ style E fill:#ffecec,stroke:#333,stroke-width:2px,color:#111
88
+ style F fill:#edf9ed,stroke:#333,stroke-width:2px,color:#111
89
+ style G fill:#fff4dd,stroke:#333,stroke-width:2px,color:#111
90
+ ```
91
+
92
+ ---
93
+
94
+ **To view or edit the diagram:** The `ako_architecture.mmd` file can be opened in any [Mermaid](https://mermaid.js.org/) editor (e.g., [Mermaid Live Editor](https://mermaid.live/)), or viewed directly on platforms like GitHub that support `.mmd` rendering.
docs/readmes/Free-Wiggum-README.md ADDED
@@ -0,0 +1,310 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Free-Wiggum
2
+ **Repository:** Julien-ser/Free-Wiggum
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** A free implementation of the ever intriguing Ralph Wiggum loops
6
+ **Language:** JavaScript
7
+ **Last Updated:** 2026-03-12T03:23:50Z
8
+
9
+ ---
10
+
11
+
12
+
13
+ # 🤖 Autonomous AI Task Loop with OpenCode
14
+
15
+ Run autonomous AI task loops **completely free.** No subscriptions. No API bills. No infrastructure costs.
16
+
17
+ While CrewAI, AutoGPT, and Claude API cost **$200-500+/month**, this setup uses [OpenCode](https://github.com/ripienaar/opencode) with OpenRouter's free models to build autonomous agents on your own machine for **$0**.
18
+
19
+ This is a lightweight, recursive bash-driven loop that uses free LLM access to work through task lists methodically until completion.
20
+
21
+ ## 🛠 Prerequisites
22
+
23
+ - **OpenCode AI** (`npm i -g opencode-ai`)
24
+ - **GitHub CLI** (`gh`) — Required for authentication bypass (avoids broken OAuth handshake)
25
+ - **OpenRouter API Key** (Free tier available at [openrouter.ai](https://openrouter.ai))
26
+ - **Linux/macOS/Windows with bash** (Git Bash on Windows works fine)
27
+
28
+ ## 📝 Example Project
29
+
30
+ **This repository itself is a real example of a project built using the Wiggum loop!**
31
+
32
+ The [Free-Wiggum](.) folder contains a fully functional **Todo Application** that was built entirely by the autonomous AI agent:
33
+ - **Express.js backend** with SQLite database
34
+ - **Responsive HTML/CSS/JS frontend** with real-time task management
35
+ - **Full test suite** for both backend and database operations
36
+ - **All development tracked** in [TASKS.md](TASKS.md) (17 completed tasks)
37
+
38
+ Look at [TASKS.md](TASKS.md) to see how the agent tracked every feature from project setup to deployment. This demonstrates exactly how the loop works in practice.
39
+
40
+ Run it locally:
41
+ ```bash
42
+ npm install
43
+ npm start
44
+ # Visit http://localhost:3000
45
+ ```
46
+
47
+ ## 📋 System Setup
48
+
49
+ ### Step 1: Install OpenCode AI
50
+
51
+ ```bash
52
+ npm i -g opencode-ai
53
+ ```
54
+
55
+ ### Step 2: Authenticate with GitHub CLI
56
+
57
+ Use the GitHub CLI to establish a persistent session that OpenCode borrows:
58
+
59
+ ```bash
60
+ gh auth login
61
+ ```
62
+
63
+ Choose **Device Code** flow when prompted. This avoids the broken OAuth browser handshake and creates a system-level authentication that OpenCode automatically uses.
64
+
65
+ ### Step 3: Set OpenRouter API Key & Token Limits
66
+
67
+ Add your OpenRouter API key and token management settings to `.env` in your project directory:
68
+
69
+ ```
70
+ OPENROUTER_API_KEY=sk-or-v1-...
71
+ WIGGUM_MODEL=openrouter/stepfun/step-3.5-flash:free
72
+
73
+ # Pre-emptive Token Management (Critical for long-running loops)
74
+ # Forces OpenCode to truncate responses at 32k tokens, leaving 32k for input context (total 64k budget)
75
+ OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=32000
76
+ ```
77
+
78
+ **Token Management Explained:**
79
+ - **OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=32000** — Prevents context bloat by capping each OpenCode response at 32k tokens
80
+ - Leaves 32k tokens for input prompts (tasks, context, AGENTS.md)
81
+ - Total budget: 64k tokens per iteration (the Wiggum rule)
82
+ - When limit is reached, the loop automatically resets and continues fresh
83
+ - Prevents model degradation after long sessions
84
+
85
+ The `wiggum.sh` script loads this automatically at startup.
86
+
87
+ ## 💰 Cost Comparison
88
+
89
+ | Tool | Monthly Cost | Setup | Model Quality |
90
+ |------|-------------|-------|---------------|
91
+ | **This Setup** | **$0** | 5 minutes | Excellent (Gemini, Qwen, Step) |
92
+ | Claude API | $20-600+ | Easy | Excellent |
93
+ | CrewAI + GPT-4 | $400+ | Medium | Excellent |
94
+ | AutoGPT | $200+ | Complex | Good |
95
+ | AWS SageMaker Agents | $500+ | Hard | Good |
96
+
97
+ **This setup uses OpenRouter's free tier**, which provides legitimate, powerful models:
98
+ - **Google Gemini 2.0 Flash** (best overall)
99
+ - **Qwen 3** (excellent for reasoning)
100
+ - **Step 3.5 Flash** (reliable and lightweight)
101
+
102
+ 1000 agent iterations = **$0.00**. Same iterations on Claude API = **$300+**.
103
+ ## 🚀 Quick Start
104
+
105
+ ### 1. Install Dependencies
106
+
107
+ ```bash
108
+ # Install OpenCode AI globally
109
+ npm i -g opencode-ai
110
+
111
+ # Authenticate with GitHub (use Device Code flow)
112
+ gh auth login
113
+ ```
114
+
115
+ ### 2. Create Your Project Structure
116
+
117
+ Create these three core files in your project directory. Use **[TASKStemplate.md](TASKStemplate.md)** as your template for TASKS.md:
118
+
119
+ **A. TASKS.md** (The Memory)
120
+
121
+ This tracks the agent's own state. The agent reads this, completes one task at a time, and marks it as done `[x]`. Start with the structure in [TASKStemplate.md](TASKStemplate.md):
122
+
123
+ ```markdown
124
+ # Current Tasks
125
+
126
+ - [ ] Task 1: Your instruction here
127
+ - [ ] Task 2: Your instruction here
128
+ - [ ] Task 3: Your instruction here
129
+ - [ ] MISSION ACCOMPLISHED
130
+ ```
131
+
132
+ **Important:** The script stops when it finds `[x] MISSION ACCOMPLISHED`. Until the agent marks this, the loop continues.
133
+
134
+ **B. AGENTS.md** (The Brain)
135
+
136
+ This file is auto-generated by OpenCode and contains project context the agent uses to understand your codebase architecture. Create it by running:
137
+
138
+ ```bash
139
+ opencode .
140
+ ```
141
+
142
+ Then inside the OpenCode interface, type:
143
+
144
+ ```
145
+ /init
146
+ ```
147
+
148
+ This generates `AGENTS.md` which the loop references via `@AGENTS.md` syntax.
149
+
150
+ **C. prompt.txt** (The Logic)
151
+
152
+ This is the system instruction sent to the agent on every iteration. It defines the agent's behavior and tells it what to do.
153
+
154
+ **C. wiggum.sh** (The Engine)
155
+
156
+ The bash script that drives the autonomous loop. Copy this repo's `wiggum.sh` and make it executable.
157
+
158
+ ### 3. Create .env File
159
+
160
+ ```bash
161
+ cat > .env << 'EOF'
162
+ OPENROUTER_API_KEY=sk-or-v1-your_key_here
163
+ WIGGUM_MODEL=openrouter/stepfun/step-3.5-flash:free
164
+ EOF
165
+ ```
166
+
167
+ ### 4. Launch the Loop
168
+
169
+ ## 📍 How It Works
170
+
171
+ 1. **GitHub Authentication**: Uses persistent `gh cli` session instead of OAuth2 (avoids browser handshake failures)
172
+ 2. **Initialization**: Loads `.env` with OpenRouter API key and model selection
173
+ 3. **Project Context**: Reads `@AGENTS.md` to understand your codebase structure
174
+ 4. **Iteration Loop**:
175
+ - Extracts the FIRST incomplete task (marked with `- [ ]`) from TASKS.md
176
+ - Builds a dynamic prompt containing:
177
+ - System instructions from `prompt.txt`
178
+ - Project context from `@AGENTS.md`
179
+ - Current `TASKS.md` state
180
+ - The specific task to complete
181
+ - Saves the complete iteration session to `logs/iteration-{iteration}.md` with formatted output for debugging
182
+ - Runs `opencode --message "$prompt" --enable-search --yes`:
183
+ - `--message`: Sends the task directly (non-interactive mode)
184
+ - `--enable-search`: Enables DuckDuckGo/MCP web search for live documentation
185
+ - `--yes`: Auto-approves file writes and shell command execution
186
+ - Waits for OpenCode to complete and update TASKS.md
187
+ 5. **Completion**: Checks if `[x] MISSION ACCOMPLISHED` is marked; if so, exits gracefully
188
+
189
+ ## 🎯 Choosing a Model
190
+
191
+ The script defaults to `openrouter/stepfun/step-3.5-flash:free`. You can change this in `wiggum.sh` at the line with `--model`.
192
+
193
+ To avoid the "yapping" loop (where the model talks but doesn't code), use models with high instruction-following and action-taking capabilities.
194
+
195
+ **Recommended Free/Cheap Models:**
196
+
197
+ - `openrouter/google/gemini-2.0-flash-exp:free` - Best overall performance
198
+ - `openrouter/qwen/qwen-3-next-80b-a3b-instruct` - Good for complex logic
199
+ - `openrouter/arcee-ai/trinity-mini` - Good for agent tasks
200
+ - `openrouter/stepfun/step-3.5-flash:free` - Lightweight and reliable
201
+
202
+ ## 🔍 Token & Context Management
203
+
204
+ ### The 64k Rule
205
+
206
+ The Wiggum loop operates under a **64k token budget per iteration**:
207
+
208
+ - **32k tokens** — Maximum output from OpenCode (via `OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=32000`)
209
+ - **32k tokens** — Reserved for input context (tasks, AGENTS.md, prompt.txt)
210
+ - **Why:** Models degrade significantly after 64k tokens in a single session
211
+
212
+ ### How It Works
213
+
214
+ 1. **Pre-emptive Capping** — `.env` variable forces OpenCode to truncate responses at 32k
215
+ 2. **Soft Guardrails** — Dynamic prompt includes explicit token constraints warning
216
+ 3. **Active Monitoring** — Script checks log file size and estimates tokens
217
+ 4. **Error Detection** — Loop detects OpenRouter "context_length_exceeded" errors and resets
218
+ 5. **Auto-Reset** — When 64k threshold is hit, token counter resets for next iteration
219
+
220
+ ### What Happens at the Limit
221
+
222
+ - Loop warns: `⚠️ Token limit (64000) reached!`
223
+ - Resets token counter: `🔄 Resetting token counter for next iteration cycle...`
224
+ - Continues with fresh iteration (no context loss, just fresh session)
225
+ - `TASKS.md` tracks all progress, so agent knows what's been done
226
+
227
+ This prevents the "yapping loop" symptom where models start repeating commands or forgetting the original goal.
228
+
229
+ OpenCode includes powerful built-in capabilities, automatically enabled via the `wiggum.sh` loop:
230
+
231
+ ### GitHub Integration (via `gh` CLI)
232
+
233
+ The authentication bridge is already established when you run `gh auth login`:
234
+
235
+ ```bash
236
+ gh auth login
237
+ # Choose: "GitHub.com" → "HTTPS" → "Yes to use git credential" → "Device Code"
238
+ ```
239
+
240
+ Benefits for your agent:
241
+ - Automatically reads your GitHub context
242
+ - Can push commits to your repositories
243
+ - Access to GitHub Issues and PRs
244
+ - No need for manual Personal Access Tokens
245
+
246
+ ### Code Search (via `--enable-search`)
247
+
248
+ The `wiggum.sh` script automatically enables search with the `--enable-search` flag:
249
+
250
+ ```bash
251
+ opencode --message "$prompt" --enable-search --yes
252
+ ```
253
+
254
+ This gives OpenCode:
255
+ - **DuckDuckGo web search** via MCP (free, no API key needed)
256
+ - **Codebase semantic search** to find existing implementations
257
+ - **Real-time documentation access** for libraries and tools
258
+
259
+ The loop uses this to:
260
+ - Find implementation patterns without external paid APIs
261
+ - Avoid code duplication by discovering similar functions
262
+ - Reference external libraries and best practices during execution
263
+
264
+ ### Configuring Search in Your Agent
265
+
266
+ Add these capabilities to your `prompt.txt`:
267
+
268
+ ```
269
+ You have access to:
270
+ - Web search: Use the search tool to find relevant documentation and examples
271
+ - Code search: Find existing functions before implementing new ones
272
+ - GitHub context: Your repositories and issues are available
273
+ - Execute freely: Use shell commands and write files without asking permission
274
+ ```
275
+
276
+ ## 🔍 Debugging
277
+
278
+ **Iteration Logs**: All runs are logged to `logs/iteration-{iteration}.md` with formatted markdown containing:
279
+ - The exact prompt sent
280
+ - OpenCode's complete output
281
+ - TASKS.md state after the iteration
282
+
283
+ Review these if a task fails.
284
+
285
+ **AGENTS.md Issues**: If tasks seem to lack project understanding, regenerate context:
286
+ ```bash
287
+ opencode .
288
+ # Inside interface, type: /reinit
289
+ ```
290
+
291
+ **OpenCode Not Executing**: If the AI doesn't write files or run commands, ensure:
292
+ - You're using `--yes` flag (auto-approves execution)
293
+ - `.env` has valid `OPENROUTER_API_KEY`
294
+ - `gh auth login` was successful
295
+
296
+ **Search Not Working**:
297
+ - Verify `--enable-search` is in the `wiggum.sh` opencode call
298
+ - Check that `OPENROUTER_API_KEY` is set (search requires API access)
299
+
300
+ **Loop Stuck**: Check `logs/iteration-*.md` to see what prompt was sent and what OpenCode responded with.
301
+
302
+ ## ⚠️ Known Issues
303
+
304
+ - **OAuth Handshake Failures**: Avoided by using `gh auth login` with Device Code flow instead of browser-based OAuth
305
+ - **Login Walls**: Models may hang on login screens; use `prompt.txt` to instruct the agent to skip authentication
306
+ - **Stuck Loops**: If no tasks complete, check `prompt-*.md` logs to see what context the agent received
307
+ - **Windows Compatibility**: Use Git Bash or WSL for full bash script support
308
+ - **Timeout Issues**: If OpenCode takes too long, increase the `sleep 3` duration in `wiggum.sh`
309
+
310
+
docs/readmes/Git-Workshop-README.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Git-Workshop
2
+ **Repository:** Julien-ser/Git-Workshop
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** JavaScript
7
+ **Last Updated:** 2026-02-21T04:50:49Z
8
+
9
+ ---
10
+
11
+ # Git Workshop Network
12
+
13
+ A visual network graph showing contributors and their connections. Perfect for your first open source contribution!
14
+
15
+ ## Quick Start
16
+
17
+ 1. Open `src/index.html` in your browser
18
+ 2. Click on nodes or sidebar cards to see contributor details
19
+ 3. Zoom and pan to explore the network
20
+
21
+ ## Contributing
22
+
23
+ See [CONTRIBUTING.md](CONTRIBUTING.md) for how to add yourself to the network.
docs/readmes/Our-Papers-README.md ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Our-Papers
2
+ **Repository:** CyberScienceLab/Our-Papers
3
+ **Contribution:** 🤝 Contributor
4
+ **Stars:** 2
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2025-11-13T15:12:49Z
8
+
9
+ ---
10
+
11
+ # Our-Papers
12
+
13
+ Here you may find our papers codes in this repo. The name of the director is the paper title and you can find the paper citation in the Readme file in the same folder!
14
+
15
+ Note: all codes in this repo are GNU v.3 license. Please cite the relevant paper when using our codes.
docs/readmes/ResumeWorthy-README.md CHANGED
@@ -4,7 +4,7 @@
4
  **Stars:** 0
5
  **Description:** Allows your resume to be split into components, easier to modularize and customize to jobs
6
  **Language:** TypeScript
7
- **Last Updated:** 2025-12-29T21:59:16Z
8
 
9
  ---
10
 
 
4
  **Stars:** 0
5
  **Description:** Allows your resume to be split into components, easier to modularize and customize to jobs
6
  **Language:** TypeScript
7
+ **Last Updated:** 2026-03-01T22:49:44Z
8
 
9
  ---
10
 
docs/readmes/Stockmetricagent-README.md ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Stockmetricagent
2
+ **Repository:** Julien-ser/Stockmetricagent
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-02-16T18:18:45Z
8
+
9
+ ---
10
+
11
+ # Agentic Stock Dashboard
12
+
13
+ A professional Streamlit app powered by LangChain agents and OpenRouter API. Users chat with an intelligent agent to generate stock insights, visualizations, and sentiment analysis in real-time. Modular, extensible, and ready for advanced financial workflows.
14
+
15
+ ## Features
16
+
17
+ - 💬 **Interactive Chat Interface** - Real-time Q&A for stock queries
18
+ - 🤖 **LangChain Agents** - Intelligent agentic workflows for complex analysis
19
+ - 📊 **Dynamic Dashboards** - Beautiful visualizations powered by Plotly
20
+ - 😊 **Sentiment Analysis** - Real-time sentiment scoring using VADER
21
+ - 🔗 **API Integration** - YahooQuery for stock data, OpenRouter for LLM
22
+ - 🏗️ **Modular Design** - Easily extensible tools and utilities
23
+
24
+ ## Prerequisites
25
+
26
+ - Python 3.10+
27
+ - OpenRouter API key (for LLM access)
28
+ - Virtual environment (stocks venv included)
29
+
30
+ ## Installation
31
+
32
+ 1. **Activate the virtual environment:**
33
+ ```powershell
34
+ .\stocks\Scripts\activate
35
+ ```
36
+
37
+ 2. **Create a `.env` file** in the project root with your API keys:
38
+ ```
39
+ OPENROUTER_API_KEY=your_api_key_here
40
+ ```
41
+
42
+ 3. **Dependencies are pre-installed** in the venv. Core packages include:
43
+ - `langchain` - Agent orchestration
44
+ - `streamlit` - Web interface
45
+ - `yahooquery` - Stock data
46
+ - `plotly` - Visualizations
47
+ - `vaderSentiment` - Sentiment analysis
48
+ - `requests` - HTTP client
49
+
50
+ ## Running the App
51
+
52
+ ```powershell
53
+ streamlit run app.py
54
+ ```
55
+
56
+ The app will launch at `http://localhost:8501`
57
+
58
+ ## Project Structure
59
+
60
+ ```
61
+ agentic_stock_dashboard/
62
+ ├── app.py # Main Streamlit application
63
+ ├── agents.py # LangChain agent configuration
64
+ ├── tools.py # Custom tools for stock analysis
65
+ ├── utils.py # Helper functions and utilities
66
+ ├── README.md # This file
67
+ └── .gitignore # Git ignore rules
68
+ ```
69
+
70
+ ### Key Files
71
+
72
+ - **app.py**: Main Streamlit entry point. Handles chat UI and agent orchestration
73
+ - **agents.py**: LangChain agent setup with tool bindings
74
+ - **tools.py**: Custom tools for stock analysis, data retrieval, and dashboard generation
75
+ - **utils.py**: Utility functions for data processing and formatting
76
+
77
+ ## Usage
78
+
79
+ 1. Start the app with `streamlit run app.py`
80
+ 2. Enter queries about stocks, sectors, or market insights
81
+ 3. The agent will analyze your query and return:
82
+ - Text-based insights and analysis
83
+ - Interactive dashboard visualizations
84
+ - Sentiment analysis results
85
+
86
+ ## Agents & Tools
87
+
88
+ ### Core Agent: `stock_agent(query)`
89
+ The main LangChain agent that orchestrates all tools. It:
90
+ 1. Parses user queries to extract stock symbols and analysis type
91
+ 2. Suggests relevant stocks for sector queries using LLM
92
+ 3. Fetches data and generates comprehensive insights
93
+ 4. Returns formatted analysis with visualizations
94
+
95
+ ### Tools Used by the Agent
96
+
97
+ #### 1. **get_stock_metrics(symbol)**
98
+ Fetches fundamental and valuation metrics for a stock using YahooQuery API.
99
+
100
+ **Returns:**
101
+ - Price, Market Cap, Enterprise Value
102
+ - PE Ratios (Trailing & Forward)
103
+ - Valuation ratios (Price/Sales, Price/Book, EV/Revenue, EV/Earnings)
104
+ - Profitability margins (Profit Margin, Operating Margin)
105
+ - Financial data (Revenue, Gross Profit, Debt, Debt/Equity)
106
+ - Ownership metrics (Insider %, Institution %, Payout Ratio, Dividend Yield)
107
+
108
+ #### 2. **get_sector_top_stocks(sector)**
109
+ Retrieves top stocks within a given sector. Used to identify leading companies in industry sectors when users query broad categories.
110
+
111
+ **Usage:** When user asks about sectors (e.g., "AI stocks", "Tech companies")
112
+
113
+ #### 3. **get_sentiment(symbol)**
114
+ Performs sentiment analysis on stock symbols using VADER (Valence Aware Dictionary and sEntiment Reasoner) from vaderSentiment library.
115
+
116
+ **Returns:** Sentiment scores indicating positive/negative/neutral sentiment around a stock
117
+
118
+ #### 4. **get_deepseek_insight(query)**
119
+ Placeholder for AI-powered insights (currently returns template response). Can be extended to call external insight APIs.
120
+
121
+ #### 5. **plot_stock_dashboard(stocks_data)**
122
+ Creates interactive visualizations for stock data using Plotly.
123
+
124
+ **Visualizations Generated:**
125
+ - **Key Metrics Table** - Formatted display of all stock metrics with appropriate units ($, %, decimals)
126
+ - **Valuation Radar Chart** - 5-axis radar chart comparing:
127
+ - Dividend Yield (ideal: 5%+)
128
+ - Operating Margin (ideal: 15-25%)
129
+ - PE Ratio (ideal: 15-25)
130
+ - Profit Margin (ideal: 15-25%)
131
+ - Institution Ownership (ideal: 33-34%)
132
+
133
+ **Chart Features:**
134
+ - Mobile-optimized (350px height, responsive margins)
135
+ - Color-coded for visual appeal
136
+ - Real-world metric ranges for accurate scoring
137
+
138
+ ## Example Queries
139
+
140
+ - "What are today's top performing stocks?"
141
+ - "Give me a technical analysis of AAPL"
142
+ - "What's the sentiment around Tesla stock?"
143
+ - "Compare NVIDIA vs AMD performance"
144
+
145
+ ## Environment Variables
146
+
147
+ Create a `.env` file with:
148
+ - `OPENROUTER_API_KEY`: Your OpenRouter API key (required)
149
+
150
+ **Note:** The `.env` file is in `.gitignore` to protect your API keys
151
+
152
+ ## License
153
+
154
+ MIT License - Feel free to use and modify as needed
docs/readmes/WiggumLoopAgenticSWDeveloper-README.md ADDED
@@ -0,0 +1,526 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # WiggumLoopAgenticSWDeveloper
2
+ **Repository:** Julien-ser/WiggumLoopAgenticSWDeveloper
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Shell
7
+ **Last Updated:** 2026-03-27T23:02:31Z
8
+
9
+ ---
10
+
11
+ # 🐳 WiggumLoopAgenticSWDeveloper
12
+
13
+ > **Zero-cost, fully autonomous multi-project AI development.** Orchestrate unlimited OpenCode agents running free LLMs to build, test, and ship software—without writing a single line of code yourself.
14
+
15
+
16
+ <!-- UI Overview -->
17
+ ![Wiggum Full Interface Overview (ui3.png)](./ui3.png)
18
+ <p align="center"><i>WiggumLoop: Full system UI overview (ui3.png)</i></p>
19
+
20
+ <!-- Optionally keep the old screenshot for reference -->
21
+ <!-- ![Wiggum UI](./ui.png) -->
22
+ ---
23
+
24
+ ## 🖼️ User Interface
25
+
26
+ WiggumLoop comes with a visual dashboard for monitoring and controlling projects. Below are detailed UI screenshots:
27
+
28
+ <p align="center">
29
+ <img src="./ui1.png" alt="Project UI Detail 1" width="48%"/>
30
+ <img src="./ui2.png" alt="Project UI Detail 2" width="48%"/>
31
+ </p>
32
+ <p align="center"><i>Detailed project views: left (ui1.png), right (ui2.png)</i></p>
33
+
34
+ The main overview (see top) shows the entire orchestration system, while these detailed views focus on individual project management and status.
35
+
36
+ ---
37
+
38
+ ---
39
+
40
+ ## ⚡ TL;DR
41
+
42
+ What if you could spawn **10+ autonomous AI developers** that each own a project, write code, run tests, and push to GitHub—**for free**?
43
+
44
+ That's WiggumLoop:
45
+
46
+ 1. **Create a project** → `bash wiggum_master.sh create "my-api" "Build a REST API"`
47
+ 2. **Write tasks** in `projects/my-api/TASKS.md`
48
+ 3. **Start the master loop** → `bash wiggum.sh`
49
+ 4. **Watch it work** → Agents code, commit, push, repeat
50
+
51
+ No babysitting. No API bills. Just autonomous development.
52
+
53
+ ---
54
+
55
+ ## 🤖 What Is This?
56
+
57
+ WiggumLoopAgenticSWDeveloper is a **master orchestration system** that:
58
+
59
+ | ✅ What it Does | 🎯 Why It Matters |
60
+ |----------------|-------------------|
61
+ | Manages **multiple projects** simultaneously | Each project gets its own GitHub repo and autonomous agent |
62
+ | Spawns **independent AI agents** using OpenCode | Projects develop in parallel, 24/7 |
63
+ | Orchestrates **master + worker loops** | System-level tasks create/coordinate projects |
64
+ | Includes **web dashboard + voice control** (optional) | Monitor and control from a browser |
65
+ | Costs **$0/month** (OpenRouter free tier) | Unlimited iterations, no API fatigue |
66
+ | Runs **Gemini 2.0 Flash, Qwen 3, Step 3.5 Flash** | Top-tier free models, not toy LLMs |
67
+
68
+ **Each project is a separate GitHub repository** that develops independently while the master system coordinates everything.
69
+
70
+ ---
71
+
72
+ ## 🏗️ System Architecture
73
+
74
+ ```mermaid
75
+ graph TB
76
+ subgraph "Master Control"
77
+ Master["🎛️ Master<br/>(wiggum_master.sh)"]
78
+ MasterLoop["🔄 Master Loop<br/>(wiggum.sh)"]
79
+ end
80
+
81
+ subgraph "Projects & Agents"
82
+ P1["📦 Project A<br/>projects/project-a/"]
83
+ P2["📦 Project B<br/>projects/project-b/"]
84
+ P3["📦 Project C<br/>projects/project-c/"]
85
+
86
+ A1["🤖 Agent A<br/>OpenCode Loop"]
87
+ A2["🤖 Agent B<br/>OpenCode Loop"]
88
+ A3["🤖 Agent C<br/>OpenCode Loop"]
89
+ end
90
+
91
+ subgraph "External Services"
92
+ GitHub["🔗 GitHub<br/>Multiple Repos"]
93
+ API["🌐 OpenRouter<br/>Free LLMs"]
94
+ end
95
+
96
+ Master -->|creates| P1
97
+ Master -->|creates| P2
98
+ Master -->|creates| P3
99
+
100
+ MasterLoop -->|spawns| A1
101
+ MasterLoop -->|spawns| A2
102
+ MasterLoop -->|spawns| A3
103
+
104
+ P1 -->|runs in| A1
105
+ P2 -->|runs in| A2
106
+ P3 -->|runs in| A3
107
+
108
+ A1 -->|commits to| GitHub
109
+ A2 -->|commits to| GitHub
110
+ A3 -->|commits to| GitHub
111
+
112
+ A1 -->|queries| API
113
+ A2 -->|queries| API
114
+ A3 -->|queries| API
115
+ ```
116
+
117
+ ---
118
+
119
+ ## 🔄 Detailed Workflow
120
+
121
+ ### Master Loop Flow
122
+
123
+ ```mermaid
124
+ flowchart TD
125
+ Start["Start wiggum.sh"] --> ReadSys["Read system TASKS.md"]
126
+ ReadSys --> CheckNew["Check for new projects?"]
127
+
128
+ CheckNew -->|Yes| CreateProj["Create project via<br/>wiggum_master.sh"]
129
+ CreateProj --> InitGit["Initialize GitHub repo"]
130
+ InitGit --> SpawnWorkers
131
+
132
+ CheckNew -->|No| SpawnWorkers
133
+
134
+ SpawnWorkers["Spawn worker agents<br/>for each projects/*/"] --> AgentRun["Each agent:<br/>- Read TASKS.md<br/>- Complete next task<br/>- Commit & push<br/>- Mark complete"]
135
+
136
+ AgentRun --> CheckComplete["All agents<br/>MISSION ACCOMPLISHED?"]
137
+
138
+ CheckComplete -->|No| Sleep["Sleep interval"]
139
+ Sleep --> Start
140
+
141
+ CheckComplete -->|Yes| Done["✅ System complete"]
142
+ ```
143
+
144
+ ### Agent Iteration Loop
145
+
146
+ ```mermaid
147
+ flowchart LR
148
+ subgraph "Agent Loop (per project)"
149
+ A1["Read TASKS.md"] --> A2["Extract first<br/> unchecked task"]
150
+ A2 --> A3["Build prompt:<br/>prompt.txt + TASKS.md + task"]
151
+ A3 --> A4["Run OpenCode<br/>(free LLM)"]
152
+ A4 --> A5["Write code<br/>Commit & push"]
153
+ A5 --> A6["Mark task [x]"]
154
+ A6 --> A7["Check MISSION ACCOMPLISHED?"]
155
+ A7 -->|No| A1
156
+ A7 -->|Yes| Done["✅ Project complete"]
157
+ end
158
+ ```
159
+
160
+ ---
161
+
162
+ ## 💰 Why This Is a Game-Changer
163
+
164
+ | | **WiggumLoop** | **Traditional AI Coding** |
165
+ |---|---|---|
166
+ | **Cost** | $0 (free LLMs) | $20–600+/month |
167
+ | **Setup Time** | 10 minutes | 30 min–2 hours |
168
+ | **Projects** | Unlimited (one repo each) | Usually one project per API key |
169
+ | **Autonomy** | 24/7 worker loops | Manual prompting per task |
170
+ | **Scalability** | Add more projects, no extra cost | Linear cost increase |
171
+
172
+ **You could run this for a year and still pay nothing.** The only thing it costs is your time to set up and guide it.
173
+
174
+ ---
175
+
176
+ ## 🚀 Quick Start (5 minutes)
177
+
178
+ ### 1️⃣ Install Prerequisites
179
+
180
+ ```bash
181
+ # Node.js (for OpenCode)
182
+ curl -fsSL https://deb.nodesource.com/setup_18.x | sudo -E bash -
183
+ sudo apt-get install -y nodejs
184
+
185
+ # OpenCode AI (global install)
186
+ npm install -g opencode-ai
187
+
188
+ # GitHub CLI (auth)
189
+ gh auth login --web
190
+
191
+ # Python deps (for optional web dashboard)
192
+ pip install -r requirements.txt
193
+ ```
194
+
195
+ ### 2️⃣ Configure
196
+
197
+ ```bash
198
+ # Get free OpenRouter key: https://openrouter.ai
199
+ cat > .env << 'EOF'
200
+ OPENROUTER_API_KEY=sk-or-v1-YOUR_KEY_HERE
201
+ WIGGUM_MODEL=openrouter/google/gemini-2.0-flash-exp:free
202
+ EOF
203
+
204
+ # Initialize system context for OpenCode
205
+ opencode /init --yes
206
+ ```
207
+
208
+ ### 3️⃣ Create Your First Project
209
+
210
+ ```bash
211
+ bash wiggum_master.sh create "todo-api" "Build a FastAPI REST API with SQLite and CRUD operations"
212
+ ```
213
+
214
+ This creates `projects/todo-api/` with its own GitHub repo.
215
+
216
+ ### 4️⃣ Add Tasks
217
+
218
+ Edit `projects/todo-api/TASKS.md`:
219
+
220
+ ````markdown
221
+ ## Todo API
222
+
223
+ - [ ] Set up FastAPI app with CORS
224
+ - [ ] Create SQLite models (Item, User)
225
+ - [ ] Implement POST /items endpoint
226
+ - [ ] Implement GET /items endpoint
227
+ - [ ] Add PATCH /items/{id} endpoint
228
+ - [ ] Add DELETE /items/{id} endpoint
229
+ - [ ] Write unit tests (pytest)
230
+ - [ ] Add request validation (Pydantic)
231
+ - [ ] Create requirements.txt
232
+ - [ ] Write README with example curl commands
233
+ - [x] MISSION ACCOMPLISHED
234
+ ````
235
+
236
+ ### 5️⃣ Start the Master Loop
237
+
238
+ ```bash
239
+ bash wiggum.sh
240
+ ```
241
+
242
+ That's it. The master will:
243
+ - Read system `TASKS.md`
244
+ - Create new projects if needed
245
+ - Spawn a worker agent for `todo-api`
246
+ - Agent reads `projects/todo-api/TASKS.md`, completes first task, commits, pushes, repeats
247
+
248
+ ---
249
+
250
+ ## 📂 Repository Structure
251
+
252
+ ```
253
+ WiggumLoopAgenticSWDeveloper/
254
+ ├── wiggum.sh # 🚀 Master orchestrator (run this)
255
+ ├── wiggum_master.sh # Project management CLI
256
+ ├── server.py # Flask web dashboard
257
+ ├── voice_server.py # Voice control (optional)
258
+ ├── prompt.txt # Master agent instructions
259
+ ├── TASKS.md # System-level tasks
260
+ ├── AGENTS.md # Auto-generated system context
261
+ ├── .env # Configuration
262
+ ├── ui.png # Dashboard screenshot
263
+
264
+ ├── project_template/ # Scaffold for new projects
265
+ │ ├── README.md
266
+ │ ├── TASKS.md
267
+ │ ├── prompt.txt
268
+ │ └── src/
269
+
270
+ ├── projects/ # Managed projects (separate GitHub repos)
271
+ │ ├── todo-api/
272
+ │ │ ├── TASKS.md
273
+ │ │ ├── prompt.txt
274
+ │ │ ├── logs/
275
+ │ │ └── src/
276
+ │ └── another-project/
277
+
278
+ └── agents/ # Specialized agent roles
279
+ ├── generic.md
280
+ ├── devops-engineer.md
281
+ └── ...
282
+ ```
283
+
284
+ ---
285
+
286
+ ## 🎯 Creating & Managing Projects
287
+
288
+ ### Create a Project
289
+
290
+ ```bash
291
+ bash wiggum_master.sh create "project-name" "What this project should do"
292
+ ```
293
+
294
+ **What happens:**
295
+ - Creates `projects/project-name/`
296
+ - Copies `project_template/` files
297
+ - Initializes Git repo
298
+ - Creates GitHub remote: `github.com/YOU/project-name`
299
+ - Pushes initial commit
300
+
301
+ ### Add Tasks
302
+
303
+ Edit `projects/project-name/TASKS.md`:
304
+
305
+ ```markdown
306
+ ## Project Name
307
+
308
+ - [ ] First task (agent does this first)
309
+ - [ ] Second task
310
+ - [ ] Third task
311
+ - [x] MISSION ACCOMPLISHED
312
+ ```
313
+
314
+ **Pro tip:** Break tasks into small, verifiable chunks. Agents work better with clear, atomic objectives.
315
+
316
+ ### Control Projects
317
+
318
+ ```bash
319
+ bash wiggum_master.sh list # List all projects
320
+ bash wiggum_master.sh status # Show status (running/stopped)
321
+ bash wiggum_master.sh stop "name" # Stop a worker
322
+ bash wiggum_master.sh start "name" # Start a worker
323
+ ```
324
+
325
+ ---
326
+
327
+ ## 🎭 Agent Roles: Specialize Your Workers
328
+
329
+ Each project can switch between specialized roles for different phases.
330
+
331
+ | Role | Best For |
332
+ |------|----------|
333
+ | `generic` | General full-stack development |
334
+ | `devops-engineer` | CI/CD, deployment, infra |
335
+ | `qa-specialist` | Testing, quality gates |
336
+ | `release-manager` | Versioning, releases |
337
+ | `documentation-specialist` | Docs, READMEs, API specs |
338
+ | `project-orchestrator` | Planning, delegation |
339
+
340
+ **Switch roles:**
341
+
342
+ ```bash
343
+ cd projects/my-project
344
+ echo "devops-engineer" > .agent_role
345
+ git add .agent_role && git commit -m "ops: switch to devops-engineer" && git push
346
+ ```
347
+
348
+ Next iteration, the agent loads specialized instructions from `agents/devops-engineer.md`.
349
+
350
+ ---
351
+
352
+ ## 🔁 Resilience: Stuck Detection & Recovery
353
+
354
+ The worker **automatically detects** when a task is stuck (no progress for 5 iterations) and applies recovery strategies:
355
+
356
+ 1. **Decompose** – breaks the task into subtasks
357
+ 2. **Skeleton files** – creates minimal structure to unblock
358
+ 3. **Skip & retry later** – moves on, will retry later
359
+
360
+ You don't have to micromanage. The system self-corrects.
361
+
362
+ ---
363
+
364
+ ## 🚨 CI/CD Error Handling
365
+
366
+ When builds/tests fail, the worker extracts the error and:
367
+
368
+ - **Code errors?** → Agent fixes the code
369
+ - **Dependency/version errors?** → Agent updates version constraints
370
+ - **Environment setup errors?** → Mark as `[CI-SKIP]`, document as prerequisite
371
+
372
+ It **never** installs system tools or downloads large files. It only modifies code, configs, and version numbers.
373
+
374
+ ---
375
+
376
+ ## 🛠️ Advanced Features
377
+
378
+ ### Web Dashboard (Optional)
379
+
380
+ ```bash
381
+ pip install -r requirements.txt
382
+ python3 server.py
383
+ # Visit: http://localhost:5000
384
+ ```
385
+
386
+ Features:
387
+ - Real-time project status
388
+ - Create/stop/start projects from UI
389
+ - View logs inline
390
+ - Trigger CodeRabbit reviews
391
+
392
+ ### Background Operation
393
+
394
+ ```bash
395
+ # Run master loop in background
396
+ nohup bash wiggum.sh > logs/master.log 2>&1 &
397
+ ```
398
+
399
+ ### Logs
400
+
401
+ Each project logs iterations to `projects/<name>/logs/`. Inspect with:
402
+
403
+ ```bash
404
+ tail -f projects/todo-api/logs/iteration-*.log
405
+ ```
406
+
407
+ ---
408
+
409
+ ## ⚙️ Configuration
410
+
411
+ ### `.env` Options
412
+
413
+ ```bash
414
+ OPENROUTER_API_KEY=sk-or-v1-... # Required (free from openrouter.ai)
415
+ WIGGUM_MODEL=openrouter/google/gemini-2.0-flash-exp:free
416
+ GITHUB_USER=your-username
417
+ MASTER_SLEEP_INTERVAL=300 # Seconds between master loops
418
+ ```
419
+
420
+ ### Recommended Free Models
421
+
422
+ - `openrouter/google/gemini-2.0-flash-exp:free` ⭐ Best overall
423
+ - `openrouter/qwen/qwen-3-80b` – Strong reasoning
424
+ - `openrouter/stepfun/step-3.5-flash:free` – Fast & reliable
425
+
426
+ ---
427
+
428
+ ## 🐛 Troubleshooting
429
+
430
+ | Problem | Fix |
431
+ |---------|-----|
432
+ | `opencode: command not found` | `npm install -g opencode-ai` |
433
+ | GitHub auth fails | `gh auth logout && gh auth login --web` |
434
+ | API key invalid | Get new key at openrouter.ai, update `.env` |
435
+ | Agent loops hanging | Check `logs/master.log`, kill process, verify `wiggum.sh` is running |
436
+ | Project repo not created | `gh auth status` – ensure you're logged in |
437
+
438
+ ---
439
+
440
+ ## 📚 Resources
441
+
442
+ - [OpenCode GitHub](https://github.com/ripienaar/opencode)
443
+ - [OpenRouter Docs](https://openrouter.ai/docs)
444
+ - [GitHub CLI Manual](https://cli.github.com/manual)
445
+ - [Free Model List](https://openrouter.ai/models)
446
+
447
+ ---
448
+
449
+ ## 📋 Core System Files
450
+
451
+ ### `wiggum.sh` — The Engine
452
+
453
+ The master orchestrator. It:
454
+ - Reads `TASKS.md` for system tasks
455
+ - Creates projects via `wiggum_master.sh`
456
+ - Spawns worker agents for each project
457
+ - Sleeps and repeats
458
+
459
+ **This is the only script you need to run** for full autonomy.
460
+
461
+ ### `wiggum_master.sh` — Project Manager
462
+
463
+ CLI for creating, listing, starting, stopping projects. Handles GitHub repo creation and project scaffolding.
464
+
465
+ ### `prompt.txt` — Agent Instructions
466
+
467
+ The system prompt sent to OpenCode on every iteration. Defines agent behavior, constraints, and workflow.
468
+
469
+ ### `TASKS.md` — Task Tracking
470
+
471
+ Markdown checklist. The master loop reads this to know what to do. Each project has its own `TASKS.md`. The loop stops when it finds `[x] MISSION ACCOMPLISHED`.
472
+
473
+ ### `project_template/` — Blueprint
474
+
475
+ When you create a project, it's copied from here. Customize this template to change default project structure.
476
+
477
+ ---
478
+
479
+ ## 🎬 Example: From Zero to Autonomous Project
480
+
481
+ ```bash
482
+ # 1. Create project
483
+ bash wiggum_master.sh create "weather-bot" "Telegram bot that posts daily forecast"
484
+
485
+ # 2. Add tasks (edit projects/weather-bot/TASKS.md)
486
+ # - [ ] Set up python-telegram-bot
487
+ # - [ ] Integrate OpenWeatherMap API
488
+ # - [ ] Schedule daily message at 8 AM
489
+ # - [ ] Add error handling + logging
490
+ # - [x] MISSION ACCOMPLISHED
491
+
492
+ # 3. Start master loop
493
+ bash wiggum.sh
494
+
495
+ # 4. Done. Watch GitHub repo get commits.
496
+ ```
497
+
498
+ The agent will:
499
+ - Build a Telegram bot skeleton
500
+ - Add weather API integration
501
+ - Implement scheduling
502
+ - Add error handling
503
+ - Mark each task complete as it goes
504
+ - Push to GitHub
505
+
506
+ ---
507
+
508
+ ## 🦄 Why This Is One of My Favorite Projects
509
+
510
+ - **It just works.** Set it and forget it. Agents run for days without intervention.
511
+ - **Zero cost = zero guilt.** Run as many experiments as you want. Fail fast, learn faster.
512
+ - **You stay in control.** Tasks are plain Markdown. No proprietary UI lock-in.
513
+ - **It scales linearly.** Want 5 more projects? Just create them. No extra API cost.
514
+ - **It's transparent.** Every prompt is saved as `prompt-*.md`. Every commit is on GitHub.
515
+
516
+ This isn't just a coding assistant—it's an **autonomous development team** that never sleeps, never asks for a raise, and never bills you by the hour.
517
+
518
+ ---
519
+
520
+ ## 📖 License
521
+
522
+ MIT. Do whatever you want with it.
523
+
524
+ ---
525
+
526
+ **Built with OpenCode. Powered by free LLMs. Orchestrated by Wiggum.**
docs/readmes/agentic-founders-finding-README.md ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # agentic-founders-finding
2
+ **Repository:** Julien-ser/agentic-founders-finding
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-12T17:42:10Z
8
+
9
+ ---
10
+
11
+ # Founder Finding Agent
12
+
13
+ An AI-powered agent that searches for startup founders across multiple platforms (X/Twitter, LinkedIn, Y Combinator, DuckDuckGo) based on a specified domain or sector.
14
+
15
+ ## Features
16
+
17
+ - **Multi-platform search**: Query X (Twitter), LinkedIn, Y Combinator, and DuckDuckGo simultaneously
18
+ - **AI-powered extraction**: Uses OpenAI or Anthropic LLMs to extract structured founder information
19
+ - **Deduplication**: Fuzzy matching to consolidate duplicate founder records
20
+ - **Scoring & ranking**: Ranks founders by relevance, social presence, and recency
21
+ - **Flexible output**: Export results as JSON or CSV
22
+
23
+ ## Prerequisites
24
+
25
+ - Python 3.11+ (if running locally)
26
+ - Docker & Docker Compose (recommended for containerized deployment)
27
+ - API keys for at least one LLM provider (OpenAI or Anthropic)
28
+ - X (Twitter) API credentials (required for Twitter platform)
29
+
30
+ ## Quick Start (Docker)
31
+
32
+ 1. **Clone and navigate**:
33
+ ```bash
34
+ cd projects/agentic-founders-finding
35
+ ```
36
+
37
+ 2. **Create `.env` file** from template:
38
+ ```bash
39
+ cp .env.example .env
40
+ ```
41
+
42
+ Edit `.env` and add your API keys (see [Configuration](#configuration) below).
43
+
44
+ 3. **Run with Docker Compose**:
45
+ ```bash
46
+ # Search for AI startup founders
47
+ docker-compose run --rm founder-finder "AI startups" --max-results 20
48
+
49
+ # Save results to file
50
+ docker-compose run --rm founder-finder "fintech" --max-results 30 --output results.json
51
+
52
+ # Use Anthropic instead of OpenAI
53
+ docker-compose run --rm founder-finder "biotech" --model-provider anthropic
54
+ ```
55
+
56
+ 4. **View output**: Results are saved to `./output/` directory (mounted volume).
57
+
58
+ ## Local Setup (Without Docker)
59
+
60
+ 1. **Install dependencies**:
61
+ ```bash
62
+ pip install -r requirements.txt
63
+ ```
64
+
65
+ 2. **Create `.env` file**:
66
+ ```bash
67
+ cp .env.example .env
68
+ ```
69
+ Edit `.env` with your API credentials.
70
+
71
+ 3. **Run the agent**:
72
+ ```bash
73
+ python -m src.main "AI startups" --max-results 10
74
+
75
+ # With output file
76
+ python -m src.main "cleantech" --max-results 25 --output founders.csv --format csv
77
+ ```
78
+
79
+ ## Configuration
80
+
81
+ Create a `.env` file with the following variables:
82
+
83
+ ### Required
84
+ - `OPENAI_API_KEY` or `ANTHROPIC_API_KEY` - LLM provider API key
85
+ - `TWITTER_API_KEY` - X (Twitter) API key
86
+ - `TWITTER_API_SECRET` - X API secret
87
+ - `TWITTER_ACCESS_TOKEN` - X access token
88
+ - `TWITTER_ACCESS_SECRET` - X access secret
89
+
90
+ ### Optional
91
+ - `LINKEDIN_USERNAME` - LinkedIn email (optional, may not work due to API restrictions)
92
+ - `LINKEDIN_PASSWORD` - LinkedIn password (optional)
93
+ - `MODEL_PROVIDER` - "openai" or "anthropic" (default: openai)
94
+ - `MODEL_NAME` - Specific model to use (optional, defaults to provider's default)
95
+ - `DUCKDUCKGO_TIMEOUT` - Search timeout in seconds (default: 30)
96
+ - `MAX_RESULTS_PER_PLATFORM` - Maximum results per platform (default: 50)
97
+
98
+ ## Command-line Options
99
+
100
+ ```bash
101
+ usage: src.main [-h] [--max-results MAX_RESULTS] [--output OUTPUT]
102
+ [--format {json,csv}] [--model-provider {openai,anthropic}]
103
+ domain
104
+
105
+ positional arguments:
106
+ domain Domain/sector to search for (e.g., 'AI', 'fintech', 'biotech')
107
+
108
+ optional arguments:
109
+ -h, --help show this help message and exit
110
+ --max-results MAX_RESULTS, -m MAX_RESULTS
111
+ Maximum results per platform (default: 10, max: 100)
112
+ --output OUTPUT, -o OUTPUT
113
+ Output file path (if not specified, prints to stdout)
114
+ --format {json,csv} Output format (default: json)
115
+ --model-provider {openai,anthropic}
116
+ LLM provider to use for extraction (default: openai)
117
+ ```
118
+
119
+ ## Output Format
120
+
121
+ Each founder record includes:
122
+
123
+ ```json
124
+ {
125
+ "name": "John Doe",
126
+ "company": "Startup Inc",
127
+ "domain": "AI",
128
+ "location": "San Francisco, CA",
129
+ "source": "twitter",
130
+ "confidence": 0.95,
131
+ "score": 0.87,
132
+ "bio_snippet": "Founder @StartupInc building AI products...",
133
+ "url": "https://twitter.com/johndoe"
134
+ }
135
+ ```
136
+
137
+ ## Project Structure
138
+
139
+ ```
140
+ agentic-founders-finding/
141
+ ├── src/
142
+ │ ├── agent/ # LangChain agent components
143
+ │ │ ├── extractor.py # LLM-based founder extraction
144
+ │ │ ├── prompt.py # System prompts
145
+ │ │ └── tools.py # Platform tool definitions
146
+ │ ├── platforms/ # Platform integrations
147
+ │ │ ├── duckduckgo.py
148
+ │ │ ├── linkedin.py
149
+ │ │ ├── twitter.py
150
+ │ │ └── yc.py
151
+ │ ├── output/ # Export functionality
152
+ │ ├── utils/ # Deduplication and scoring
153
+ │ └── main.py # CLI entry point
154
+ ├── tests/ # Unit and integration tests
155
+ ├── docker-compose.yml # Docker Compose configuration
156
+ ├── Dockerfile # Docker image build
157
+ ├── requirements.txt # Python dependencies
158
+ └── .env.example # Environment variable template
159
+ ```
160
+
161
+ ## Docker Details
162
+
163
+ - **Multi-stage build**: Uses Python 3.11-slim with non-root user for security
164
+ - **Lightweight**: Builder stage installs dependencies, runtime stage is minimal
165
+ - **Persistent storage**: Output directory mounted as volume (`./output`)
166
+ - **Configuration**: Environment variables loaded from host `.env` file (mounted read-only)
167
+
168
+ ### Build and Run Manually
169
+
170
+ ```bash
171
+ # Build image
172
+ docker build -t founder-finder .
173
+
174
+ # Run
175
+ docker run --rm --env-file .env -v $(pwd)/output:/app/output founder-finder "AI startups"
176
+ ```
177
+
178
+ ## Testing
179
+
180
+ Run tests locally:
181
+
182
+ ```bash
183
+ pytest tests/ -v
184
+ ```
185
+
186
+ Run with Docker:
187
+
188
+ ```bash
189
+ docker-compose run --rm founder-finder pytest tests/ -v
190
+ ```
191
+
192
+ ## Notes
193
+
194
+ - LinkedIn API access may be limited; the `linkedin-api` package may not work reliably with modern LinkedIn
195
+ - X (Twitter) API requires a paid tier for full search capabilities
196
+ - Results quality depends on available API rate limits and query specificity
197
+ - For production use, consider adding persistent storage (database) and async processing
198
+
199
+ ## License
200
+
201
+ MIT
docs/readmes/agentic-slide-filler-README.md ADDED
@@ -0,0 +1,148 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # agentic-slide-filler
2
+ **Repository:** Julien-ser/agentic-slide-filler
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-29T03:00:58Z
8
+
9
+ ---
10
+
11
+ # Agentic Slide Filler
12
+
13
+ An AI-powered tool that automatically generates professional PowerPoint presentations from document outlines using LLMs.
14
+
15
+ ## Architecture
16
+
17
+ ### Technology Stack
18
+ - **Language**: Python 3.11+
19
+ - **PPT Manipulation**: `python-pptx` (>=0.6.21)
20
+ - **AI Generation**: OpenAI API (GPT-4/3.5) or Anthropic API
21
+ - **Document Parsing**: `python-docx` (>=0.8.11), Markdown support
22
+ - **CLI**: argparse
23
+ - **Testing**: pytest
24
+
25
+ ### System Architecture
26
+
27
+ ```
28
+ ┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
29
+ │ Input Files │ │ AI Content │ │ Output File │
30
+ │ • Template │───▶│ Generation │───▶│ • presentation │
31
+ │ • Outline │ │ │ │ .pptx │
32
+ └─────────────────┘ └──────────────────┘ └─────────────────┘
33
+ │ │ │
34
+ ▼ ▼ ▲
35
+ ┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
36
+ │ TemplateParser │ │ AIContent │ │ PPTXWriter │
37
+ │ • Layout extract│ │ Generator │ │ • Populate │
38
+ │ • Placeholders │ │ • API calls │ │ • Formatting │
39
+ └─────────────────┘ │ • Rate limiting │ └─────────────────┘
40
+ │ │ • Caching │ │
41
+ ▼ └──────────────────┘ │
42
+ ┌─────────────────┘ ▲ │
43
+ │ OutlineParser │ │ │
44
+ │ • Docx/MD parse │───────────────┘ │
45
+ │ • Hierarcy JSON │ │
46
+ └─────────────────┘ │
47
+ │ │
48
+ └────────────────┬─────────────────────────────┘
49
+
50
+ ┌─────────────────────┐
51
+ │ ContentMapper │
52
+ │ • Section align │
53
+ │ • Placeholder match │
54
+ └─────────────────────┘
55
+ ```
56
+
57
+ ### Data Flow
58
+ 1. **Parse Template**: Extract slide layouts, placeholders, and shape types
59
+ 2. **Parse Outline**: Convert document (DOCX/MD) to hierarchical content structure
60
+ 3. **Map Content**: Align outline sections with appropriate template placeholders
61
+ 4. **Generate AI Content**: Call LLM API with prompts to generate slide content
62
+ 5. **Validate Content**: Check length, formatting, and compatibility
63
+ 6. **Write PPTX**: Populate template with generated content and save
64
+
65
+ ## Setup
66
+
67
+ ```bash
68
+ # Install dependencies
69
+ pip install python-pptx openai python-docx pytest PyYAML python-dotenv
70
+
71
+ # Or with requirements.txt
72
+ pip install -r requirements.txt
73
+
74
+ # Set up environment
75
+ cp .env.example .env
76
+ # Edit .env and add your OPENAI_API_KEY
77
+
78
+ # Optional: customize configuration
79
+ # Copy and modify config.yaml as needed
80
+ cp config.yaml config.local.yaml
81
+
82
+ # Run the tool
83
+ python -m src.cli --template template.pptx --outline outline.docx --output presentation.pptx
84
+ ```
85
+
86
+ ## Configuration
87
+
88
+ The application supports configuration through:
89
+
90
+ **Environment Variables** (`.env` file):
91
+ - `OPENAI_API_KEY` - Your OpenAI API key (required)
92
+ - `ANTHROPIC_API_KEY` - Anthropic API key (optional, for Claude)
93
+
94
+ **Configuration File** (`config.yaml` or `config.local.yaml`):
95
+ - OpenAI model settings (model name, temperature, max tokens)
96
+ - File paths (templates, output, cache, logs)
97
+ - Output filename patterns
98
+ - Content generation parameters (retry attempts, caching)
99
+ - Logging configuration
100
+
101
+ Configuration precedence: environment variables > `config.local.yaml` > `config.yaml`
102
+
103
+ ## Project Structure
104
+
105
+ ```
106
+ .
107
+ ├── README.md # This file
108
+ ├── TASKS.md # Development progress
109
+ ├── requirements.txt # Python dependencies
110
+ ├── config.yaml # Application configuration
111
+ ├── config.local.yaml # Local overrides (optional)
112
+ ├── .env.example # Environment variable template
113
+ ├── .env # Local environment variables (gitignored)
114
+ ├── src/ # Source code
115
+ │ ├── template_parser.py
116
+ │ ├── outline_parser.py
117
+ │ ├── content_mapper.py
118
+ │ ├── ai_generator.py
119
+ │ ├── validator.py
120
+ │ ├── pptx_writer.py
121
+ │ ├── slide_filler.py
122
+ │ ├── prompts.py
123
+ │ └── cli.py
124
+ ├── tests/ # Test suite
125
+ ├── templates/ # Sample PPT templates
126
+ ├── output/ # Generated presentations
127
+ ├── cache/ # LLM response cache
128
+ └── logs/ # Application logs
129
+ ```
130
+
131
+ ## Current Status
132
+
133
+ ✅ **Phase 1**: Planning & Setup - Complete
134
+ 🔄 **Phase 2**: Template & Outline Processing - In Progress
135
+
136
+ ### Completed
137
+ - ✅ Template parser (`src/template_parser.py`): Extracts slide layouts, placeholders, and metadata from PPTX templates
138
+ - ✅ Outline parser (`src/outline_parser.py`): Extracts hierarchical content from DOCX and Markdown documents
139
+ - ✅ Content mapper (`src/content_mapper.py`): Aligns outline sections with template placeholders using intelligent matching
140
+ - ✅ Comprehensive unit tests (`tests/test_template_parser.py`) covering placeholder detection, layout identification, and error handling
141
+ - ✅ Comprehensive unit tests (`tests/test_outline_parser.py`) covering DOCX/Markdown parsing, hierarchy extraction, and error handling
142
+
143
+ ### In Progress
144
+ - 🔄 Adding template validation
145
+
146
+ ### Up Next
147
+ - Design LLM prompt templates (`src/prompts.py`)
148
+ - Build AI content generator (`src/ai_generator.py`)
docs/readmes/agentic-team-README.md ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # agentic-team
2
+ **Repository:** Julien-ser/agentic-team
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-14T02:24:45Z
8
+
9
+ ---
10
+
11
+ # Agentic Team
12
+
13
+ An autonomous multi-agent development system built on the wiggum loop concept. Three specialized agents (Security, Software Developer, Frontend) collaborate via Agent-to-Agent (A2A) communication to complete development tasks autonomously.
14
+
15
+ ## Architecture Overview
16
+
17
+ The system consists of:
18
+ - **3 Specialized Agents**: Security, Software Developer, Frontend
19
+ - **Redis Message Broker**: A2A communication backbone
20
+ - **SQLite Shared State**: Task persistence and coordination
21
+ - **Flask Dashboard**: Real-time monitoring and control
22
+ - **Wiggum Master Loop**: Task dispatcher and orchestrator
23
+
24
+ Read the full architecture documentation in [`docs/architecture.md`](docs/architecture.md).
25
+
26
+ ## Current Progress
27
+
28
+ **Phase 1 - Planning & Architecture** (Complete)
29
+ - ✅ Task 1.1: System architecture and component diagram completed
30
+ - ✅ Task 1.2: Define agent role specifications and protocols
31
+ - ✅ Task 1.3: Create database schema for shared state
32
+ - ✅ Task 1.4: Setup project dependencies and environment configuration
33
+
34
+ **Phase 2 - Core Infrastructure & Wiggum Loop Enhancement** (Complete)
35
+ - ✅ Task 2.1: Implement the enhanced wiggum loop with role-based agent selection
36
+ - ✅ Task 2.2: Build the message broker using Redis pub/sub
37
+ - ✅ Task 2.3: Create agent base class and lifecycle manager
38
+ - ✅ Task 2.4: Implement shared state manager with SQLite
39
+
40
+ **Phase 3 - Specialized Agent Workers** (Complete)
41
+ - ✅ Task 3.1: Implement Security Agent with vulnerability scanning & code review
42
+ - ✅ OWASP Top 10 2021 compliance validation integrated (`src/security/owasp_validator.py`)
43
+ - ✅ Comprehensive security pattern detection (secrets, SQL injection, XSS, SSRF, etc.)
44
+ - ✅ Dependency CVE auditing with safety/pip-audit
45
+ - ✅ Automated security alerts and recommendations to other agents
46
+ - ✅ Task 3.2: Implement Software Development Agent for backend code generation
47
+ - ✅ Task 3.3: Implement Frontend Agent for UI/UX development
48
+ - ✅ Task 3.4: Build agent worker orchestration with health monitoring
49
+
50
+ **Phase 4 - A2A Communication & Integration Testing** (Complete)
51
+ - ✅ Task 4.1: Implement A2A message routing and handling
52
+ - ✅ Task 4.2: Build collaborative workflow: end-to-end feature development
53
+ - ✅ Task 4.3: Create web dashboard for monitoring agent activity
54
+ - ✅ Task 4.4: Write comprehensive documentation and finalize TASKS.md
55
+
56
+ ## Quick Start
57
+
58
+ ### Prerequisites
59
+ - Python 3.12+
60
+ - Redis (running on localhost:6379)
61
+ - SQLite (included with Python)
62
+
63
+ ### Setup
64
+
65
+ 1. Install dependencies:
66
+ ```bash
67
+ pip install -r requirements.txt
68
+ ```
69
+
70
+ 2. Configure environment (optional):
71
+ ```bash
72
+ cp .env.example .env
73
+ # Edit .env with your settings
74
+ ```
75
+
76
+ 3. Initialize the database:
77
+ ```bash
78
+ python -m src.state.migrate
79
+ ```
80
+
81
+ 4. Start the system:
82
+ ```bash
83
+ python -m src.orchestrator.main
84
+ ```
85
+
86
+ 5. Launch the dashboard (separate terminal):
87
+ ```bash
88
+ python -m src.dashboard.app
89
+ ```
90
+
91
+ 6. Visit `http://localhost:5000` to monitor agent activity
92
+
93
+ ## Project Structure
94
+
95
+ ```
96
+ agentic-team/
97
+ ├── docs/ # Documentation
98
+ │ ├── architecture.md # System design and specifications
99
+ │ └── workflow_example.md # Collaborative workflow walkthrough
100
+ ├── src/ # Source code
101
+ │ ├── protocols/ # Agent specifications and message protocols
102
+ │ ├── agents/ # Specialized agent implementations
103
+ │ ├── messaging/ # Redis broker and message router
104
+ │ ├── state/ # SQLite state manager
105
+ │ ├── core/ # Wiggum loop implementation
106
+ │ ├── orchestrator/ # Agent worker orchestration
107
+ │ └── dashboard/ # Flask monitoring interface
108
+ ├── tests/ # Unit and integration tests
109
+ │ └── test_collaborative_workflow.py # End-to-end workflow integration test
110
+ ├── requirements.txt # Python dependencies
111
+ ├── .env.example # Environment configuration template
112
+ └── TASKS.md # Development roadmap
113
+ ```
114
+
115
+ ## How It Works
116
+
117
+ 1. **Task Definition**: Tasks with role tags (`[SECURITY]`, `[SW_DEV]`, `[FRONTEND]`) are defined in `TASKS.md`
118
+ 2. **Task Dispatch**: Wiggum Master parses tasks and assigns them to appropriate agents
119
+ 3. **Collaboration**: Agents communicate via Redis pub/sub, exchanging messages and code
120
+ 4. **Shared State**: All interactions and task status persisted in SQLite
121
+ 5. **Monitoring**: Flask dashboard displays real-time agent activity and system metrics
122
+
123
+ ## Agent Capabilities
124
+
125
+ - **Security Agent**:
126
+ - OWASP Top 10 2021 compliance validation
127
+ - Vulnerability scanning (secrets, SQL injection, XSS, SSRF, etc.)
128
+ - Dependency CVE auditing (safety, pip-audit)
129
+ - Security recommendations and A2A alerts
130
+ - **Software Dev Agent**: Code generation, unit testing, formatting (black), linting (ruff), refactoring
131
+ - **Frontend Agent**: UI component generation, responsive design (Tailwind CSS), accessibility, API integration
132
+
133
+ ## Message Protocol
134
+
135
+ Agents communicate using typed messages with Pydantic validation. See [`docs/architecture.md`](docs/architecture.md#message-protocol-specification) for complete specification.
136
+
137
+ ## License
138
+
139
+ MIT
docs/readmes/basestation_server-README.md CHANGED
@@ -4,7 +4,7 @@
4
  **Stars:** 0
5
  **Description:** None
6
  **Language:** JavaScript
7
- **Last Updated:** 2025-08-09T13:00:53Z
8
 
9
  ---
10
 
 
4
  **Stars:** 0
5
  **Description:** None
6
  **Language:** JavaScript
7
+ **Last Updated:** 2026-03-04T14:01:34Z
8
 
9
  ---
10
 
docs/readmes/calorie-counter-README.md ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # calorie-counter
2
+ **Repository:** Julien-ser/calorie-counter
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** JavaScript
7
+ **Last Updated:** 2026-03-13T03:34:53Z
8
+
9
+ ---
10
+
11
+ # Calorie Counter
12
+
13
+ A full-stack web application for tracking daily calorie intake with a React frontend and Node.js/Express backend.
14
+
15
+ ## Current Status
16
+
17
+ **Phase 1 (Planning & Setup)**: ✅ Complete
18
+ **Phase 2 (Backend)**: ✅ Complete - Express server with middleware, SQLite database, REST API endpoints with calorie calculation and date-based filtering fully implemented.
19
+ **Phase 3 (Frontend)**: ✅ Complete - React app with functional components, hooks, Context API for state management, MealForm and MealList components fully implemented.
20
+ **Phase 4 (Testing & Polish)**: ✅ Complete - Backend and frontend tests implemented, responsive CSS styling with clean UI design, mobile-first approach, and enhanced user experience.
21
+
22
+ ## Project Scope
23
+
24
+ Build a calorie tracking system that allows users to:
25
+ - Log meals with food name, calorie count, date, and meal type
26
+ - View meals grouped by date with daily calorie totals
27
+ - Query a food database for calorie information
28
+ - Filter meals by date range
29
+
30
+ ## User Stories
31
+
32
+ 1. **As a user**, I want to add a meal entry with food name, calories, date, and meal type (breakfast/lunch/dinner/snack) so that I can track what I eat.
33
+
34
+ 2. **As a user**, I want to view all my meals grouped by date so that I can see my consumption patterns.
35
+
36
+ 3. **As a user**, I want to see daily calorie totals so that I can monitor my intake against goals.
37
+
38
+ 4. **As a user**, I want to delete incorrect meal entries so that my log stays accurate.
39
+
40
+ 5. **As a user**, I want to filter meals by date range so that I can focus on specific periods.
41
+
42
+ 6. **As a user**, I want quick access to common food calorie values so that I can log meals efficiently.
43
+
44
+ ## Tech Stack
45
+
46
+ - **Frontend**: React with functional components, hooks, and Context API for state management
47
+ - **Backend**: Node.js/Express with middleware (CORS, body-parser, helmet)
48
+ - **Database**: SQLite with tables for users, foods, and meals
49
+ - **API**: RESTful endpoints for CRUD operations on meals and food lookup
50
+ - **Testing**: Jest/Supertest for backend, React Testing Library for frontend
51
+
52
+ ## Features
53
+
54
+ - Meal creation with calorie tracking
55
+ - Date-based grouping and filtering
56
+ - Daily calorie summaries
57
+ - Food database integration
58
+ - Responsive CSS styling with modern UI
59
+ - Mobile-first design approach
60
+ - Smooth animations and transitions
61
+ - Enhanced accessibility features
62
+ - Clean, gradient-based color scheme
63
+ - Interactive hover effects and visual feedback
64
+
65
+ ## Setup Instructions
66
+
67
+ 1. **Clone and install dependencies**:
68
+ ```bash
69
+ cd client && npm install
70
+ cd ../server && npm install
71
+ ```
72
+
73
+ 2. **Start the backend server**:
74
+ ```bash
75
+ cd server && npm start
76
+ ```
77
+ Server runs on http://localhost:3001
78
+
79
+ 3. **Start the frontend development server**:
80
+ ```bash
81
+ cd client && npm start
82
+ ```
83
+ App runs on http://localhost:3000
84
+
85
+ 4. **Run tests**:
86
+ ```bash
87
+ cd server && npm test
88
+ cd ../client && npm test
89
+ ```
90
+
91
+ ## API Endpoints
92
+
93
+ - `GET /api/meals` - Retrieve all meals (with optional date filtering)
94
+ - `POST /api/meals` - Create a new meal entry
95
+ - `DELETE /api/meals/:id` - Delete a meal
96
+ - `GET /api/foods` - Search food database for calorie information
97
+
98
+ ## Project Structure
99
+
100
+ ```
101
+ calorie-counter/
102
+ ├── client/ # React frontend
103
+ │ ├── src/
104
+ │ │ ├── components/
105
+ │ │ ├── context/
106
+ │ │ └── App.js
107
+ │ └── package.json
108
+ ├── server/ # Express backend
109
+ │ ├── src/
110
+ │ │ ├── routes/
111
+ │ │ ├── middleware/
112
+ │ │ └── database/
113
+ │ └── package.json
114
+ ├── TASKS.md # Development progress
115
+ └── README.md # This file
116
+ ```
docs/readmes/causal-github-pr-analysis-README.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # causal-github-pr-analysis
2
+ **Repository:** Julien-ser/causal-github-pr-analysis
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Jupyter Notebook
7
+ **Last Updated:** 2026-01-31T17:40:50Z
8
+
9
+ ---
10
+
11
+ # Causal GitHub PR Analysis
12
+
13
+ This project analyzes GitHub pull request data to uncover causal relationships and insights.
14
+
15
+ ## Files
16
+ - **causaltest.ipynb**: Jupyter notebook for data analysis and experimentation.
17
+ - **ghpr.csv**: Dataset containing GitHub pull request information.
18
+
19
+ ## Getting Started
20
+ 1. Open `causaltest.ipynb` in Jupyter or VS Code.
21
+ 2. Ensure you have Python and Jupyter installed.
22
+ 3. Install required packages (see notebook for details).
23
+ 4. Run the notebook cells to perform the analysis.
24
+
25
+ ## Project Purpose
26
+ The goal is to explore causal inference techniques on real-world GitHub PR data, enabling better understanding of factors influencing PR outcomes.
27
+
28
+ ## Requirements
29
+ - Python 3.x
30
+ - Jupyter Notebook
31
+ - pandas, numpy, matplotlib, and other data science libraries (see notebook)
32
+
33
+ ## Usage
34
+ - Modify or extend the notebook to try new analyses.
35
+ - Replace `ghpr.csv` with your own data if desired.
36
+
37
+ ## License
38
+ MIT License
docs/readmes/causal-model-README.md ADDED
@@ -0,0 +1,137 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # causal-model
2
+ **Repository:** Julien-ser/causal-model
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-12T19:30:15Z
8
+
9
+ ---
10
+
11
+ # Causal Model for Firmware Update Agent
12
+
13
+ A causal inference model and visualization dashboard for an agentic loop that manages firmware updates.
14
+
15
+ ## Mission
16
+
17
+ Create a sample causal model/visualization of an agentic loop for updating firmware. The project uses DoWhy for causal inference and Streamlit for interactive visualization.
18
+
19
+ ## Causal Model
20
+
21
+ ### DAG Structure
22
+
23
+ The causal graph includes these variables:
24
+ - **Update_Trigger**: Initiates update (manual, auto, emergency)
25
+ - **Network_Stability**: Network quality (0-1 score)
26
+ - **Device_Resources**: CPU, memory, storage availability
27
+ - **Device_Health**: Battery, temperature, hardware status
28
+ - **Device_Configuration**: Settings, encryption, policies
29
+ - **Firmware_Version**: Current version (older harder to update)
30
+ - **Update_Command**: Treatment - whether update command issued
31
+ - **Update_Success**: Outcome - whether update succeeded
32
+ - **Rollback_Needed**: Whether rollback was required
33
+ - **Device_Downtime**: Downtime duration during update
34
+
35
+ ### Key Relationships
36
+
37
+ - Main effect: `Update_Command` → `Update_Success`
38
+ - Moderators: `Network_Stability`, `Device_Resources`, `Device_Health` → `Update_Success`
39
+ - Confounders: `Firmware_Version`, `Device_Configuration` affect both treatment and outcome
40
+ - Consequences: `Update_Success` → `Rollback_Needed, Device_Downtime`
41
+
42
+ See `src/causal_graph.py` for complete definition.
43
+
44
+ ## Synthetic Data Generation
45
+
46
+ The project includes a comprehensive synthetic data generator (`src/data_generator.py`) that simulates realistic firmware update scenarios.
47
+
48
+ ### Features
49
+
50
+ - Generates 1000+ firmware update attempts with realistic confounding variables
51
+ - Respects causal DAG structure: treatment affects outcome, confounders influence both
52
+ - Categorical variables: `Update_Trigger` (manual/auto/emergency), `Firmware_Version` (v1.0-v3.0), `Device_Configuration` (low/medium/high)
53
+ - Continuous variables: `Network_Stability`, `Device_Resources`, `Device_Health` (0-1 scores)
54
+ - Binary outcomes: `Update_Command` (treatment), `Update_Success` (outcome), `Rollback_Needed`
55
+ - Continuous outcome: `Device_Downtime` (seconds)
56
+
57
+ ### Usage
58
+
59
+ ```python
60
+ from src.data_generator import generate_data, DataGenerator
61
+
62
+ # Quick generation
63
+ data = generate_data(n_samples=1000, random_seed=42)
64
+
65
+ # Custom configuration
66
+ config = DataGeneratorConfig(
67
+ n_samples=2000,
68
+ network_stability_mean=0.7,
69
+ treatment_effect_size=0.25
70
+ )
71
+ generator = DataGenerator(config)
72
+ data = generator.generate()
73
+
74
+ # Generate interventional data (set all devices to receive update command)
75
+ interventional_data = generator.generate_interventional_data(treatment_value=1)
76
+ ```
77
+
78
+ ### Testing
79
+
80
+ The data generator includes 17 comprehensive unit tests covering:
81
+ - Data structure and variable types
82
+ - Causal relationships (e.g., older firmware → more updates)
83
+ - Effect sizes (e.g., better network → higher success rate)
84
+ - Reproducibility with random seeds
85
+
86
+ Run tests: `pytest tests/test_data_generator.py -v`
87
+
88
+ ## Setup
89
+
90
+ ```bash
91
+ # Install dependencies
92
+ pip install -r requirements.txt
93
+
94
+ # Run the Streamlit dashboard (once implemented)
95
+ streamlit run app.py
96
+ ```
97
+
98
+ ## Project Structure
99
+
100
+ - `src/` - Source code (data generator, causal engine, agent, simulator, visualization)
101
+ - `tests/` - Unit and integration tests
102
+ - `app.py` - Streamlit dashboard
103
+ - `docker-compose.yml` - Docker deployment configuration
104
+ - `README.md` - This file
105
+
106
+ ## Progress
107
+
108
+ - ✅ **Phase 1** - Planning & Setup (complete)
109
+ - ✅ **Phase 2** - Causal Model Development (complete)
110
+ - Causal graph (`src/causal_graph.py`) implemented with 10 nodes and 12 edges
111
+ - Synthetic data generator (`src/data_generator.py`) creates realistic data respecting DAG
112
+ - Causal inference engine (`src/causal_engine.py`) built with DoWhy
113
+ - Comprehensive unit tests (`tests/test_causal.py`) passing (15 tests)
114
+ - ✅ **Phase 3** - Agentic Loop Implementation (complete)
115
+ - `Agent` class (`src/agent.py`) with perception, reasoning, action modules
116
+ - Policy function maps predicted success probability to update decision (threshold)
117
+ - `Simulator` (`src/simulator.py`) runs 100 decisions and tracks outcomes vs. baselines
118
+ - 🔄 **Phase 4** - Visualization & Dashboard (complete)
119
+ - Causal graph visualization (`src/visualization.py`) with Plotly
120
+ - Streamlit dashboard (`app.py`) with three tabs: Causal Graph, Agent Performance, Scenario Simulator
121
+ - Real-time metrics panel: agent win rate, false positives, rollback frequency
122
+ - Interactive scenario simulator
123
+ - This README and `docker-compose.yml`
124
+
125
+ ## Running the Dashboard
126
+
127
+ 1. Install dependencies: `pip install -r requirements.txt`
128
+ 2. Run: `streamlit run app.py`
129
+ 3. Open `http://localhost:8501` in your browser.
130
+
131
+ Or with Docker Compose: `docker-compose up --build` then open `http://localhost:8501`.
132
+
133
+ ## Example Scenarios
134
+
135
+ - **High-probability update**: Device with good network stability, ample resources, good health, and recent firmware → agent predicts high success probability and recommends update.
136
+ - **Low-probability update**: Device with poor network, low resources, poor health, and old firmware → agent advises against update.
137
+ - **Baseline comparison**: The dashboard shows how the agent compares to always-updating or never-updating baselines, accounting for rollback costs.
docs/readmes/cuda-optimizer-README.md ADDED
@@ -0,0 +1,231 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # cuda-optimizer
2
+ **Repository:** Julien-ser/cuda-optimizer
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-13T05:01:21Z
8
+
9
+ ---
10
+
11
+ # CUDA Optimizer for PyTorch
12
+
13
+ A specialized toolkit for optimizing PyTorch neural networks on CUDA devices. Maximize throughput, minimize memory usage, and achieve production-grade performance with minimal code changes.
14
+
15
+ ## Features
16
+
17
+ - **Custom CUDA Kernels**: Fused operations (activation+layernorm) with 20%+ speedup
18
+ - **Intelligent Memory Pool**: Caching allocator reducing fragmentation by >95%
19
+ - **Automatic Mixed Precision**: Dynamic loss scaling per layer, maintaining FP32 accuracy
20
+ - **Kernel Auto-Tuner**: Automatically finds optimal block/grid dimensions
21
+ - **Gradient Checkpointing**: 50%+ memory reduction with selective recompute
22
+ - **Tensor Parallelism**: Linear scaling across multiple GPUs with NCCL
23
+ - **Fused Optimizer**: AdamW fused kernel 30% faster than standard
24
+ - **Real-time Monitoring**: Live dashboard for GPU utilization, memory, throughput
25
+
26
+ ## Quick Start
27
+
28
+ ```bash
29
+ # Install from source (requires CUDA 11.8+)
30
+ pip install -e .
31
+
32
+ # Or use Docker with pre-configured CUDA environment
33
+ docker build -t cuda-optimizer -f Dockerfile.cuda-dev .
34
+ ```
35
+
36
+ ## Usage
37
+
38
+ ```python
39
+ import torch
40
+ from cuda_optimizer import Optimizer, profile_model
41
+
42
+ # Profile your model to identify bottlenecks
43
+ profile_model(model, input_shape=(32, 3, 224, 224))
44
+
45
+ # Apply optimizations with one line
46
+ optimized_model = Optimizer.optimize(model)
47
+
48
+ # Train with 30% less memory, 20% more throughput
49
+ criterion = torch.nn.CrossEntropyLoss()
50
+ optimizer = torch.optim.AdamW(optimized_model.parameters(), lr=1e-3)
51
+
52
+ # No code changes required - just drop-in replacement!
53
+ ```
54
+
55
+ ## Architecture
56
+
57
+ ```mermaid
58
+ graph TD
59
+ A[PyTorch Model] --> B[Input Tensors]
60
+ B --> C[Memory Pool]
61
+ C --> D[Custom CUDA Kernels]
62
+ D --> E[Fused Optimizer]
63
+ E --> F[Grad Checkpoint]
64
+ F --> G[Tensor Parallel]
65
+ G --> H[Output/Loss]
66
+
67
+ H --> I[AMP Wrapper]
68
+ I --> J[Profiler]
69
+ J --> K[Autotuner]
70
+ K --> L[Config Cache]
71
+
72
+ J --> M[Monitoring Dashboard]
73
+ M --> N[Live Metrics]
74
+
75
+ L --> D
76
+
77
+ style A fill:#f9f,stroke:#333
78
+ style H fill:#bbf,stroke:#333
79
+ style N fill:#9f9,stroke:#333
80
+ ```
81
+
82
+ ## Performance Targets
83
+
84
+ | Model | FPS Improvement | Memory Reduction |
85
+ |-------|----------------|------------------|
86
+ | ResNet50 | +30% | -33% |
87
+ | BERT-small | +30% | -39% |
88
+ | LSTM | +20% | -50% |
89
+ | GPT-2 small | +25% | -50% |
90
+
91
+ Full targets and validation criteria: [docs/optimization_targets.md](docs/optimization_targets.md)
92
+
93
+ ## Requirements
94
+
95
+ - **CUDA**: 11.8+
96
+ - **cuDNN**: 8.x+
97
+ - **PyTorch**: 2.0+
98
+ - **Python**: 3.9+
99
+ - **GPU**: NVIDIA (Compute capability >= 7.0)
100
+
101
+ ## Project Structure
102
+
103
+ ```
104
+ cuda-optimizer/
105
+ ├── src/
106
+ │ ├── kernels/ # Custom CUDA kernels
107
+ │ ├── memory/ # Caching allocator
108
+ │ ├── optim/ # AMP wrapper
109
+ │ ├── tuner/ # Auto-tuning
110
+ │ ├── checkpoint/ # Gradient checkpointing
111
+ │ ├── parallel/ # Tensor parallelism
112
+ │ ├── fusion/ # Fused optimizers
113
+ │ ├── monitoring/ # Dashboard
114
+ │ └── profiling/ # Profiling tools
115
+ ├── docs/ # Documentation
116
+ ├── tests/ # Unit & integration tests
117
+ ├── scripts/ # CLI utilities
118
+ ├── notebooks/ # Tutorials 📚
119
+ │ ├── 01_basics.ipynb # CNN/ResNet50 optimization
120
+ │ ├── 02_transformers.ipynb # BERT optimization
121
+ │ └── 03_distributed.ipynb # Multi-GPU training
122
+ ├── data/ # Benchmark datasets
123
+ └── dashboard/ # Monitoring UI
124
+ ```
125
+
126
+ ## Installation
127
+
128
+ ### From Source
129
+
130
+ ```bash
131
+ # Clone and install
132
+ git clone <repo-url>
133
+ cd cuda-optimizer
134
+ pip install -e .
135
+ ```
136
+
137
+ ### Docker (Recommended for quick start)
138
+
139
+ ```bash
140
+ docker build -t cuda-optimizer -f Dockerfile.cuda-dev .
141
+ docker run --gpus all -it cuda-optimizer
142
+ ```
143
+
144
+ ## Development
145
+
146
+ ```bash
147
+ # Run baseline benchmarks
148
+ python scripts/run_baseline.py --models resnet50 bert-small
149
+
150
+ # Run tests (with GPU)
151
+ pytest tests/ --gpu
152
+
153
+ # Run tests with coverage
154
+ pytest tests/ --cov=cuda_optimizer --cov-report=html
155
+
156
+ # Run tests in parallel (requires pytest-xdist)
157
+ pytest -n auto tests/
158
+
159
+ # Run specific test suite
160
+ pytest tests/unit/ # Unit tests only
161
+ pytest tests/integration/ # Integration tests only
162
+
163
+ # Lint and type check
164
+ black src/ tests/
165
+ isort src/ tests/
166
+ flake8 src/ tests/
167
+ mypy src/ tests/
168
+
169
+ # Start monitoring dashboard
170
+ streamlit run dashboard/app.py
171
+ ```
172
+
173
+ ## Documentation
174
+
175
+ - [Quickstart Guide](docs/quickstart.md)
176
+ - [API Reference](docs/api/)
177
+ - [Migration Guide](docs/migration_guide.md)
178
+ - [Optimization Targets](docs/optimization_targets.md)
179
+ - [Troubleshooting](docs/troubleshooting.md)
180
+
181
+ ## Current Status
182
+
183
+ **Phase 1: Planning & Setup**
184
+ - ✅ Task 1.1: Define optimization targets and requirements
185
+ - ✅ Task 1.2: Set up development environment with CUDA toolchain
186
+ - ✅ Task 1.3: Establish baseline profiling infrastructure
187
+ - ✅ Task 1.4: Create project structure and dependency management
188
+
189
+ **Phase 2: Core CUDA Optimization Implementation**
190
+ - ✅ Task 2.1: Implement custom CUDA kernels for tensor operations ([learn more](docs/custom_ops.md))
191
+ - ✅ Task 2.2: Develop memory pool and caching allocator ([learn more](docs/cuda_cache.md))
192
+ - ✅ Task 2.3: Create automatic mixed precision optimizer wrapper ([learn more](docs/amp_wrapper.md))
193
+ - ✅ Task 2.4: Build kernel auto-tuner using NVIDIA NVTX ([learn more](docs/autotuner.md))
194
+
195
+ **Phase 3: Advanced Features & Integration**
196
+ - ✅ Task 3.1: Implement gradient checkpointing with custom recompute ([learn more](docs/checkpointing.md))
197
+ - ✅ Task 3.2: Develop tensor parallelism utilities ([learn more](docs/tensor_parallel.md))
198
+ - ✅ Task 3.3: Create optimizer fusion pass (AdamW fused kernel) ([learn more](docs/adam_fused.md))
199
+ - ✅ Task 3.4: Build real-time monitoring dashboard ([learn more](docs/dashboard.md))
200
+
201
+ **Phase 4: Testing, Documentation & Deployment**
202
+ - ✅ Task 4.1: Implement comprehensive test suite ([learn more](docs/testing.md))
203
+ - ✅ Task 4.2: Create user documentation and API reference
204
+ - [Quickstart Guide](docs/quickstart.md)
205
+ - [API Reference](docs/api/)
206
+ - [Migration Guide](docs/migration_guide.md)
207
+ - [Troubleshooting](docs/troubleshooting.md)
208
+ - [Optimization Targets](docs/optimization_targets.md)
209
+ - ✅ Task 4.3: Package and publish to PyPI
210
+ - ✅ Task 4.4: Create Jupyter notebooks with tutorials
211
+ - [Notebook 1: CNN Optimization](notebooks/01_basics.ipynb) (ResNet50)
212
+ - [Notebook 2: Transformer Optimization](notebooks/02_transformers.ipynb) (BERT)
213
+ - [Notebook 3: Distributed Training](notebooks/03_distributed.ipynb) (Multi-GPU)
214
+
215
+ See [TASKS.md](TASKS.md) for complete roadmap.
216
+
217
+ ## License
218
+
219
+ MIT
220
+
221
+ ## Citation
222
+
223
+ If you use this tool in your research, please cite:
224
+
225
+ ```bibtex
226
+ @software{cuda_optimizer_2026,
227
+ title = {CUDA Optimizer for PyTorch},
228
+ year = {2026},
229
+ url = {https://github.com/your-org/cuda-optimizer}
230
+ }
231
+ ```
docs/readmes/databricks-powerbi-pipeline-README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # databricks-powerbi-pipeline
2
+ **Repository:** Julien-ser/databricks-powerbi-pipeline
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-13T15:14:17Z
8
+
9
+ ---
10
+
11
+ # Databricks Power BI Pipeline
12
+
13
+ **Production-ready end-to-end analytics solution connecting Databricks to Power BI for e-commerce analytics.**
14
+
15
+ **Status: Production Ready ✅**
16
+
17
+ ## Overview
18
+
19
+ This project provides a production-ready pipeline that transforms raw e-commerce data into actionable Power BI dashboards using Databricks as the analytics engine.
20
+
21
+ ### Problem Statement
22
+ Organizations need to analyze e-commerce data (orders, customers, products) to make data-driven decisions about inventory, marketing, and customer experience.
23
+
24
+ ### Solution
25
+ Automated ETL pipeline with:
26
+ - Bronze layer: Raw data ingestion
27
+ - Silver layer: Cleaned and validated data
28
+ - Gold layer: Business-level aggregates for reporting
29
+ - Power BI integration with DirectQuery for real-time analytics
30
+
31
+ ## Architecture
32
+
33
+ ```
34
+ Data Sources → Databricks (ETL) → Delta Tables → Power BI → Dashboards
35
+ ↓ ↓ ↓ ↓ ↓
36
+ CSV/API Notebooks Optimized DirectQuery KPIs &
37
+ (PySpark) Layer Reports
38
+ ```
39
+
40
+ ### Components
41
+
42
+ - **notebooks/** - Databricks notebooks (bronze, silver, gold transformations)
43
+ - **src/** - Python automation scripts for deployment and monitoring
44
+ - **data/** - Sample datasets and schema definitions
45
+ - **config/** - Configuration files for environments
46
+ - **tests/** - Unit and integration tests
47
+ - **docs/** - Architecture and deployment guides
48
+
49
+ ## Quick Start
50
+
51
+ ### Prerequisites
52
+ - Databricks workspace (Community Edition or paid)
53
+ - Power BI Desktop (free) or Power BI Premium
54
+ - Python 3.8+ with pip
55
+
56
+ ### Setup
57
+
58
+ ```bash
59
+ # Clone and navigate
60
+ cd databricks-powerbi-pipeline
61
+
62
+ # Install dependencies
63
+ pip install -r requirements.txt
64
+
65
+ # Configure Databricks connection
66
+ cp config/env.example.py config/env.py
67
+ # Edit config/env.py with your credentials (DATABRICKS_HOST and DATABRICKS_TOKEN)
68
+
69
+ # Optional: Generate sample data (or upload your own)
70
+ python src/generate_sample_data.py
71
+
72
+ # Run tests (unit and integration)
73
+ python -m pytest tests/
74
+
75
+ # Deploy notebooks to Databricks workspace
76
+ python src/deploy_notebooks.py
77
+
78
+ # Execute the ETL pipeline (bronze → silver → gold)
79
+ python src/monitor_pipeline.py
80
+ ```
81
+
82
+ ### Using the Pipeline
83
+
84
+ **Quickest path to production:**
85
+
86
+ 1. **Configure & Deploy**:
87
+ - Set your Databricks credentials in `config/env.py`
88
+ - Run `python src/deploy_notebooks.py` to upload notebooks to workspace
89
+ - Upload sample data to `/mnt/data/raw/` in your Databricks workspace
90
+
91
+ 2. **Execute ETL**:
92
+ - Run `python src/monitor_pipeline.py` to execute the full pipeline
93
+ - The script runs bronze → silver → gold notebooks automatically
94
+ - Monitor logs in `logs/pipeline.log` for progress
95
+
96
+ 3. **Connect Power BI**:
97
+ - Follow the [Power BI Setup Guide](docs/powerbi-setup.md)
98
+ - Connect to gold Delta tables using DirectQuery
99
+ - Import the data model and build reports
100
+
101
+ 4. **Monitor and Maintain**:
102
+ - Run `make health-check` to verify pipeline health
103
+ - Check `logs/pipeline.log` for execution details
104
+ - Use `make rollback` for disaster recovery if needed
105
+ - Set up alerts via email or Slack (configure in `config/env.py`)
106
+
107
+ ## Project Structure
108
+
109
+ ```
110
+ databricks-powerbi-pipeline/
111
+ ├── notebooks/ # Databricks ETL notebooks (Medallion Architecture)
112
+ │ ├── 01_bronze/
113
+ │ │ └── bronze_ingestion.ipynb # Raw data → Delta bronze tables
114
+ │ ├── 02_silver/
115
+ │ │ └── silver_transformation.ipynb # Cleaned data → Delta silver tables
116
+ │ └── 03_gold/
117
+ │ └── gold_aggregation.ipynb # Business aggregates → Delta gold tables
118
+ ├── src/ # Python automation scripts
119
+ │ ├── deploy_notebooks.py # Deploy notebooks to Databricks
120
+ │ ├── monitor_pipeline.py # Execute pipeline with monitoring
121
+ │ ├── generate_sample_data.py # Generate synthetic e-commerce data
122
+ │ └── utils.py # Shared utilities (logging, config, etc.)
123
+ ├── data/ # Sample datasets (CSV format)
124
+ │ ├── sample_customers.csv
125
+ │ ├── sample_products.csv
126
+ │ ├── sample_orders.csv
127
+ │ └── sample_order_items.csv
128
+ ├── config/ # Configuration files
129
+ │ ├── env.py # Environment credentials (gitignored)
130
+ │ ├── env.example.py # Template for env.py
131
+ │ └── parameters.json # Pipeline parameters (paths, catalog, etc.)
132
+ ├── tests/ # Automated tests
133
+ │ ├── unit/ # Unit tests (utils, data generation)
134
+ │ │ ├── test_utils.py
135
+ │ │ └── test_data_generation.py
136
+ │ └── integration/ # Integration tests (end-to-end)
137
+ │ └── test_integration.py
138
+ ├── docs/ # Documentation
139
+ │ ├── architecture.md # System design and architecture
140
+ │ ├── deployment.md # Step-by-step deployment guide
141
+ │ └── powerbi-setup.md # Power BI connection and dashboard setup
142
+ ├── logs/ # Pipeline execution logs (auto-generated)
143
+ ├── requirements.txt # Python dependencies
144
+ ├── README.md # This file
145
+ └── TASKS.md # Project roadmap and progress tracking
146
+ ```
147
+
148
+ ## Sample Data
149
+
150
+ The project includes synthetic e-commerce data:
151
+ - **Orders**: Transaction records with timestamps, amounts, status
152
+ - **Customers**: Demographics and segmentation data
153
+ - **Products**: Catalog with categories and pricing
154
+
155
+ ## Power BI Integration
156
+
157
+ After processing data through the Delta Lake pipeline, connect Power BI for analytics. See the **[Power BI Setup Guide](docs/powerbi-setup.md)** for detailed instructions.
158
+
159
+ **Quick steps:**
160
+ 1. Get your Databricks workspace URL and token
161
+ 2. In Power BI Desktop: Get Data → Databricks
162
+ 3. Enter server and database name pointing to gold Delta table
163
+ 4. Use DirectQuery mode for real-time dashboard updates
164
+ 5. Build visuals for sales trends, customer segmentation, product performance
165
+
166
+ Or, for step-by-step instructions with screenshots, follow the complete guide in `docs/powerbi-setup.md`.
167
+
168
+ ## Testing
169
+
170
+ The project includes comprehensive tests:
171
+
172
+ ```bash
173
+ # Unit tests (fast, no external dependencies)
174
+ pytest tests/unit/
175
+
176
+ # Integration tests (run in simulation mode by default, no Databricks required)
177
+ pytest tests/integration/
178
+
179
+ # Full test suite with coverage report
180
+ pytest --cov=src tests/
181
+ ```
182
+
183
+ **Note:** Integration tests run in *simulation mode* when databricks-sdk is not installed, making them suitable for CI/CD and local development without cloud credentials. To test against a real Databricks workspace, install `databricks-sdk` and configure your credentials.
184
+
185
+ ## Development Status
186
+
187
+ See [TASKS.md](TASKS.md) for current progress and upcoming work.
188
+
189
+ ## Documentation
190
+
191
+ - [Architecture Guide](docs/architecture.md)
192
+ - [Deployment Guide](docs/deployment.md)
193
+ - [Power BI Setup](docs/powerbi-setup.md)
194
+ - [API Reference](docs/api-reference.md)
195
+ - [Configuration Reference](docs/configuration.md)
196
+ - [Monitoring & Operations](docs/monitoring.md)
197
+ - [Data Dictionary](docs/data-dictionary.md)
198
+ - [Contributing Guide](CONTRIBUTING.md)
199
+
200
+ ## Contributing
201
+
202
+ 1. Check TASKS.md for current priorities
203
+ 2. Create feature branch from main
204
+ 3. Add tests for new functionality
205
+ 4. Submit changes with clear commit messages
206
+
207
+ ## License
208
+
209
+ MIT License - see LICENSE file for details.
docs/readmes/edgebot-ai-README.md ADDED
@@ -0,0 +1,815 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # edgebot-ai
2
+ **Repository:** Julien-ser/edgebot-ai
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** C++
7
+ **Last Updated:** 2026-03-16T16:18:22Z
8
+
9
+ ---
10
+
11
+ # EdgeBot AI
12
+
13
+ [![CI](https://github.com/edgebot-ai/edgebot-ai/workflows/CI%20(Rust)/badge.svg)](https://github.com/edgebot-ai/edgebot-ai/actions)
14
+ [![Rust](https://img.shields.io/badge/rust-1.70%2B-orange.svg)](https://www.rust-lang.org)
15
+
16
+ A Rust-based platform for deploying lightweight AI models on robots and IoT devices. EdgeBot AI provides zero-copy memory safety, WebAssembly compilation, and seamless ROS2 integration for edge inference.
17
+
18
+ ## Mission
19
+
20
+ Build toolkits for deploying AI models on robots and IoT devices with:
21
+ - **Zero-copy memory safety** using Rust's ownership model
22
+ - **Cross-platform support** (x86_64, ARM, WebAssembly)
23
+ - **ROS2 integration** for robotics ecosystems
24
+ - **Burn framework** for efficient inference
25
+ - **Simulation-ready** with Webots integration
26
+
27
+ ## Project Structure
28
+
29
+ This is a Cargo workspace with multiple crates:
30
+
31
+ | Crate | Purpose | Status |
32
+ |-------|---------|--------|
33
+ | `edgebot-core` | Core inference engine + memory safety + optimizer + tasks | ✅ Phase 2 (Core SDK completed) |
34
+ | `edgebot-sim` | Simulation environment with Webots integration (headless testing) | ✅ Phase 3 (Webots integration completed) |
35
+ | `edgebot-ros2` | ROS2 bridge for robot communication | 📦 Phase 2 |
36
+ | `edgebot-wasm` | WebAssembly runtime for browser/IoT | ✅ Phase 2 (Runtime done) |
37
+ | `edgebot-sim-server` | Cloud simulation service with Actix Web API | ✅ Phase 3 |
38
+ | `edgebot-cli` | Command-line interface for compilation, deployment, simulation, and optimization | 🆕 Phase 4 (In progress) |
39
+
40
+ ## Prerequisites
41
+
42
+ - Rust 1.70+ with `rustc`, `cargo`, `rustfmt`, `clippy`
43
+ - For WASM target: `rustup target add wasm32-unknown-unknown`
44
+ - A `rust-toolchain.toml` is included to ensure consistent toolchain versions.
45
+ - For ROS2 integration: `ros2` installation (optional)
46
+
47
+ ## Setup
48
+
49
+ ```bash
50
+ # Clone and enter workspace
51
+ git clone https://github.com/edgebot-ai/edgebot-ai.git
52
+ cd edgebot-ai
53
+
54
+ # Build all crates
55
+ cargo build
56
+
57
+ # Run tests
58
+ cargo test --workspace
59
+
60
+ # Build for WASM (requires wasm32 target)
61
+ cargo build --target wasm32-unknown-unknown --release -p edgebot-wasm
62
+ ```
63
+
64
+ ## Development
65
+
66
+ ```bash
67
+ # Format code
68
+ cargo fmt
69
+
70
+ # Lint
71
+ cargo clippy --workspace -- -D warnings
72
+
73
+ # Build with optimizations
74
+ cargo build --release
75
+
76
+ # Run benchmarks (requires criterion)
77
+ cargo bench -p edgebot-core
78
+
79
+ # Generate JSON report for pro tier optimization
80
+ cargo bench -p edgebot-core -- --output-format json > benchmark_results/raw.json
81
+ # The pro tier optimization report is auto-generated at benchmark_results/inference_pro_report.json
82
+ ```
83
+
84
+ ## Current Status
85
+
86
+ **Phase 3: Simulation & Compilation** - ✅ COMPLETED
87
+
88
+ - [x] Phase 3 Task 1: Webots simulation integration
89
+ - [x] Phase 3 Task 2: Cloud simulation service
90
+ - [x] Phase 3 Task 3: ARM cross-compilation toolchain
91
+ - [x] Phase 3 Task 4: Profiling & benchmarking suite (criterion)
92
+
93
+ **Phase 4: Deployment & Monetization** - ✅ COMPLETED
94
+
95
+ - [x] Phase 4 Task 1: EdgeBot CLI (deploy, simulate, optimize commands)
96
+ - [x] Phase 4 Task 2: License verification system
97
+ - [x] Phase 4 Task 3: Dashboard frontend
98
+ - [x] Phase 4 Task 4: Comprehensive documentation
99
+
100
+ ---
101
+
102
+ ✅ **All planned development tasks are complete!** EdgeBot AI is feature-complete for initial release.
103
+
104
+ See [TASKS.md](TASKS.md) for complete roadmap.
105
+
106
+ ## EdgeBot CLI
107
+
108
+ The `edgebot-cli` is the main command-line interface for end-users, providing commands for deployment, simulation, optimization, and cross-compilation.
109
+
110
+ ### Installation
111
+
112
+ Build from the workspace:
113
+
114
+ ```bash
115
+ cargo build --release -p edgebot-cli
116
+ # Binary will be at: target/release/edgebot
117
+ ```
118
+
119
+ Or install globally:
120
+
121
+ ```bash
122
+ cargo install --path edgebot-cli
123
+ ```
124
+
125
+ ### Commands
126
+
127
+ #### 1. Compile for ARM targets
128
+
129
+ Cross-compile models for embedded ARM devices (Raspberry Pi, STM32, generic ARM).
130
+
131
+ ```bash
132
+ # Auto-detect hardware and compile
133
+ edgebot compile --model model.onnx --output binary
134
+
135
+ # Compile for specific hardware
136
+ edgebot compile --model model.onnx --output binary --hardware raspberry-pi
137
+
138
+ # Compile for all supported targets
139
+ edgebot compile --model model.onnx --output-dir ./bin/ --all
140
+
141
+ # With optimization features
142
+ edgebot compile --model model.onnx --output binary --release --features "tch,cuda"
143
+ ```
144
+
145
+ **Options:**
146
+ - `--model <path>`: Model file to embed (optional)
147
+ - `--output <path>`: Output binary path
148
+ - `--target <triple>`: Target triple (e.g., aarch64-unknown-linux-musl)
149
+ - `--hardware <type>`: Hardware type (raspberry-pi, stm32, generic-arm)
150
+ - `--all`: Build for all supported ARM targets
151
+ - `--release`: Build in release mode
152
+ - `--features <feat>`: Enable Cargo features
153
+ - `--static-link`: Statically link all dependencies
154
+
155
+ #### 2. Deploy to device
156
+
157
+ Deploy compiled binaries to remote devices via SSH or serial connection.
158
+
159
+ ```bash
160
+ # Deploy via SSH (using SSH agent)
161
+ edgebot deploy --binary ./target/release/edgebot --target 192.168.1.100 --username pi
162
+
163
+ # Deploy with password authentication
164
+ edgebot deploy --binary ./edgebot --target 192.168.1.100 --username pi --password secret
165
+
166
+ # Deploy to custom destination path
167
+ edgebot deploy --binary ./edgebot --target 192.168.1.100 --username pi --destination /home/pi/apps/
168
+
169
+ # Serial deployment (placeholder)
170
+ edgebot deploy --binary ./edgebot --target /dev/ttyUSB0 --method serial
171
+ ```
172
+
173
+ **Options:**
174
+ - `--binary <path>`: Binary file to deploy (required)
175
+ - `--target <addr>`: Target IP address or serial port (required)
176
+ - `--method <ssh|serial>`: Deployment method (default: ssh)
177
+ - `--destination <path>`: Remote path (default: /opt/edgebot/)
178
+ - `--username <user>`: SSH username (required for SSH)
179
+ - `--password <pass>`: SSH password (optional; uses SSH agent if omitted)
180
+
181
+ **Note:** Serial deployment is not yet fully implemented. Use SSH for now.
182
+
183
+ #### 3. Run simulation
184
+
185
+ Test models in either local Webots simulation or cloud simulation server.
186
+
187
+ ##### Local Simulation (Webots)
188
+
189
+ ```bash
190
+ # Run local simulation with a world file
191
+ edgebot simulate --model model.ebmodel --world worlds/warehouse.wbt --runs 10
192
+
193
+ # Output JSON results
194
+ edgebot simulate --model model.ebmodel --world worlds/warehouse.wbt --json
195
+
196
+ # Adjust simulation timestep
197
+ edgebot simulate --model model.ebmodel --world worlds/warehouse.wbt --timestep 16
198
+ ```
199
+
200
+ ##### Cloud Simulation
201
+
202
+ ```bash
203
+ # Run on cloud simulation server
204
+ edgebot simulate --model model.ebmodel --cloud --server http://localhost:8080 --runs 100
205
+
206
+ # With custom world file uploaded to server
207
+ edgebot simulate --model model.ebmodel --world worlds/custom.wbt --cloud --server http://sim.edgebot.ai
208
+ ```
209
+
210
+ **Options:**
211
+ - `--model <path>`: Model file to test (required)
212
+ - `--world <path>`: Webots world file (required for local, optional for cloud)
213
+ - `--cloud`: Use cloud simulation server
214
+ - `--server <url>`: Cloud server URL (default: http://localhost:8080)
215
+ - `--runs <n>`: Number of simulation runs (default: 1)
216
+ - `--json`: Output results as JSON
217
+ - `--timestep <ms>`: Simulation timestep in milliseconds (default: 32)
218
+
219
+ **Output:** Simulation metrics including total steps, runtime, average inference time.
220
+
221
+ #### 4. Optimize models
222
+
223
+ Optimize models for edge deployment with quantization, pruning, and layer fusion.
224
+
225
+ ```bash
226
+ # Basic optimization with int8 quantization and layer fusion
227
+ edgebot optimize --input model.onnx --output model.ebmodel --quantize int8 --fuse-layers
228
+
229
+ # Advanced: fp16 quantization + magnitude pruning (50% threshold)
230
+ edgebot optimize \
231
+ --input model.onnx \
232
+ --output model.ebmodel \
233
+ --quantize fp16 \
234
+ --prune magnitude \
235
+ --pruning-threshold 0.5 \
236
+ --device cpu
237
+
238
+ # No optimization (just convert format)
239
+ edgebot optimize --input model.onnx --output model.ebmodel --quantize none
240
+ ```
241
+
242
+ **Options:**
243
+ - `--input <path>`: Input model file (ONNX or Burn .bin) (required)
244
+ - `--output <path>`: Output optimized model (.ebmodel) (required)
245
+ - `--quantize <none|int8|fp16>`: Quantization method (default: none)
246
+ - `--prune <none|magnitude|structured>`: Pruning strategy (default: none)
247
+ - `--pruning-threshold <0.0-1.0>`: Fraction of weights to prune (default: 0.5)
248
+ - `--fuse-layers`: Enable layer fusion (Conv+ReLU, etc.)
249
+ - `--device <cpu|cuda>`: Target device (default: cpu)
250
+
251
+ **Output:** `.ebmodel` bundle containing optimized model and metadata. The CLI reports size reduction and optimization statistics.
252
+
253
+ ### Usage
254
+
255
+ **Free Tier:**
256
+ - All core SDK features are free and open source (MIT/Apache-2.0)
257
+ - Local simulation with Webots is unlimited
258
+ - Basic model optimization (quantization none, no pruning)
259
+ - Cross-compilation for ARM targets
260
+
261
+ **Pro Tier ($29/month):**
262
+ - Cloud simulation with batch runs (100+ scenarios)
263
+ - Advanced model optimization: int8/fp16 quantization, pruning, layer fusion
264
+ - Priority support and custom integrations
265
+ - Offline activation tokens available
266
+
267
+ To enable pro features, set the `EDGEBOT_LICENSE_KEY` environment variable:
268
+
269
+ ```bash
270
+ export EDGEBOT_LICENSE_KEY="your_license_key_here"
271
+ ```
272
+
273
+ The license key is an Ed25519 signed token verified offline. Contact sales@edgebot.ai to obtain a pro license.
274
+
275
+ ## EdgeBot Dashboard
276
+
277
+ The EdgeBot Dashboard is a modern web application built with Yew and Rust WebAssembly. It provides a comprehensive interface for monitoring simulation results, tracking model performance metrics, and managing your Pro subscription.
278
+
279
+ ### Features
280
+
281
+ - **Simulation Monitoring**: View real-time and historical simulation jobs, including FPS, inference latency, memory usage, and detailed performance breakdowns.
282
+ - **Model Metrics**: Track inference latency, memory footprint, and model size across different platforms (x86_64, ARM, WASM).
283
+ - **License Management**: Check your subscription status, view active features, and manage your EDGEBOT_LICENSE_KEY.
284
+
285
+ ### Building
286
+
287
+ Build the dashboard using **Trunk** (the recommended approach):
288
+
289
+ ```bash
290
+ # Install trunk (once)
291
+ cargo install trunk
292
+
293
+ # Build for production
294
+ cd edgebot-dashboard
295
+ trunk build --release
296
+ ```
297
+
298
+ Alternatively, use the included build script:
299
+
300
+ ```bash
301
+ cd edgebot-dashboard
302
+ ./build-dashboard-wasm.sh --release
303
+ ```
304
+
305
+ The compiled static files will be in the `dist/` directory.
306
+
307
+ ### Running Locally
308
+
309
+ Start a local development server with hot reloading:
310
+
311
+ ```bash
312
+ cd edgebot-dashboard
313
+ trunk serve --open
314
+ ```
315
+
316
+ Or serve the built files:
317
+
318
+ ```bash
319
+ cd edgebot-dashboard/dist
320
+ python3 -m http.server 8000
321
+ # Open http://localhost:8080
322
+ ```
323
+
324
+ ### Deployment
325
+
326
+ The dashboard is a static site and can be deployed to any static hosting service.
327
+
328
+ #### GitHub Pages
329
+
330
+ 1. Build the dashboard: `trunk build --release`
331
+ 2. Copy the build output to the `docs/` directory (which GitHub Pages uses):
332
+ ```bash
333
+ cp -r dist/* ../docs/
334
+ ```
335
+ 3. Commit and push to GitHub. GitHub Pages will automatically serve from the `docs/` folder.
336
+ Alternatively, use the provided GitHub Actions workflow (`.github/workflows/dashboard.yml`) for automatic deployment on push to `main`.
337
+
338
+ > **Note**: If your repository is served from a subpath (e.g., `https://username.github.io/edgebot-ai/`), you may need to set `public_url` in `edgebot-dashboard/trunk.toml` accordingly (e.g., `public_url = "/edgebot-ai"`).
339
+
340
+ #### Netlify
341
+
342
+ - Build command: `trunk build --release`
343
+ - Publish directory: `edgebot-dashboard/dist`
344
+ - Add an environment variable `EDGEBOT_SIM_SERVER_URL` if connecting to a remote simulation server.
345
+
346
+ ### Configuration
347
+
348
+ - **Simulation Server**: By default, the dashboard connects to `http://localhost:8080`. Override via the `EDGEBOT_SIM_SERVER_URL` environment variable.
349
+ - **License**: Pro license status is verified locally. Set `EDGEBOT_LICENSE_KEY` in the environment (or in your shell) to enable Pro features.
350
+
351
+ ### Architecture
352
+
353
+ The dashboard integrates with the EdgeBot ecosystem via two main APIs:
354
+
355
+ - **SimServerClient**: Fetches simulation jobs and results from the `edgebot-sim-server`.
356
+ - **LicensingClient**: Checks local license verification using the `edgebot-licensing` crate.
357
+
358
+ All data is fetched asynchronously using `wasm_bindgen_futures` and displayed using reactive Yew components.
359
+
360
+ ## Monetization & License Verification
361
+
362
+ EdgeBot AI uses a freemium model:
363
+
364
+ - **Free Core SDK**: All core inference, memory safety, ROS2 integration, and WebAssembly compilation remain open source under MIT/Apache-2.0.
365
+ - **Pro Features**: Cloud simulation and advanced optimization require a paid subscription.
366
+
367
+ ### License Verification System
368
+
369
+ The `edgebot-licensing` crate implements Ed25519-based license verification:
370
+
371
+ - License keys are signed by EdgeBot AI's private key
372
+ - Supports offline activation (no phone home)
373
+ - Includes expiry dates and feature flags
374
+ - Fast cryptographic verification using `ed25519-dalek`
375
+
376
+ **Command-line usage:**
377
+
378
+ ```bash
379
+ # Cloud simulation requires pro license
380
+ edgebot simulate --model model.ebmodel --cloud --server https://sim.edgebot.ai
381
+
382
+ # Model optimization requires pro license
383
+ edgebot optimize --input model.onnx --output model.ebmodel --quantize int8 --fuse-layers
384
+ ```
385
+
386
+ If no valid license is found, the commands will return an error with instructions.
387
+
388
+ **Environment Variable:**
389
+
390
+ Set `EDGEBOT_LICENSE_KEY` with your license key:
391
+
392
+ ```bash
393
+ export EDGEBOT_LICENSE_KEY="signature_base64:payload_base64"
394
+ ```
395
+
396
+ The license key format is two base64-encoded parts separated by a colon:
397
+ - `signature`: Ed25519 signature over the payload
398
+ - `payload`: JSON containing customer_id, timestamp, features, and optional expiry
399
+
400
+ **Development License:**
401
+
402
+ During development, you can generate test licenses using the `generate_dev_license` function (only available in debug builds):
403
+
404
+ ```rust
405
+ #[cfg(debug_assertions)]
406
+ let license = edgebot_licensing::generate_dev_license(
407
+ "test_customer",
408
+ vec!["cloud_sim", "optimization"],
409
+ "your_secret_key_base64"
410
+ )?;
411
+ ```
412
+
413
+ ### Obtaining a Pro License
414
+
415
+ Visit https://edgebot.ai/pricing to subscribe. After payment, you'll receive a license key via email. The key is valid for the subscription period and can be used on multiple machines.
416
+
417
+ For enterprise deployments, offline activation tokens and custom feature flags are available. Contact sales@edgebot.ai.
418
+
419
+ ### Architecture Highlights
420
+
421
+ ### Zero-Copy Memory Safety
422
+
423
+ The `edgebot-core/memory` module provides safe abstractions for sharing sensor data without memory copies between ROS2 messages and inference pipelines:
424
+
425
+ ```rust
426
+ use edgebot_core::memory::{CameraBuffer, ImageFormat, ImageMetadata, ZeroCopyBuffer};
427
+
428
+ // Create camera buffer from raw sensor data (zero-copy)
429
+ let metadata = ImageMetadata::new(640, 480, ImageFormat::RGB);
430
+ let mut buffer = CameraBuffer::new(metadata);
431
+
432
+ // Fill buffer with ROS2 image data
433
+ buffer.bytes_mut().copy_from_slice(&ros2_image_data);
434
+
435
+ // Convert to Burn tensor (minimal copy if needed)
436
+ let device = burn::backend::tch::TchBackend::Device::default();
437
+ let tensor = buffer.to_tensor(&device);
438
+
439
+ // Run inference
440
+ // let output = inference_engine.forward(tensor);
441
+ ```
442
+
443
+ **Key Types:**
444
+
445
+ - `ZeroCopyBuffer<T>`: Safe buffer using `MaybeUninit` for uninitialized memory management
446
+ - `CameraBuffer`: Zero-copy image buffer supporting RGB, BGR, RGBA, grayscale, depth
447
+ - `LidarBuffer`: Point cloud buffer for LiDAR data (xyz, intensity, ring, timestamp)
448
+ - `BorrowedBuffer<'a, T>`: Temporary view into initialized buffer data
449
+ - ROS2 integration helpers (`Ros2ImageConverter`, `Ros2PointCloudConverter`)
450
+
451
+ ### Burn Integration
452
+
453
+ The inference engine (`edgebot-core::inference`) supports multiple backends (Tch, Autocast) and model formats (ONNX, Burn binary):
454
+
455
+ ```rust
456
+ use edgebot_core::inference::InferenceEngine;
457
+ use burn::backend::tch::TchBackend;
458
+
459
+ fn main() -> Result<(), Box<dyn std::error::Error>> {
460
+ let device = TchBackend::Device::default();
461
+
462
+ // Load ONNX model
463
+ let engine = InferenceEngine::<TchBackend>::load_onnx(
464
+ Path::new("model.onnx"),
465
+ &[1, 3, 224, 224], // input shape
466
+ device,
467
+ )?;
468
+
469
+ // Or load Burn binary
470
+ // let engine = InferenceEngine::<TchBackend>::load_bin(Path::new("model.bin"), device)?;
471
+
472
+ // Create input tensor (example)
473
+ let input = burn::tensor::Tensor::<TchBackend>::random(
474
+ [1, 3, 224, 224],
475
+ burn::tensor::Distribution::Uniform(-1.0, 1.0),
476
+ engine.device(),
477
+ );
478
+
479
+ // Run inference
480
+ let output = engine.forward(input)?;
481
+ println!("Output shape: {:?}", output.dims());
482
+
483
+ Ok(())
484
+ }
485
+ ```
486
+
487
+ ### Model Optimizer
488
+
489
+ Optimize models for edge deployment with quantization, pruning, and layer fusion using the `edgebot-optimize` CLI:
490
+
491
+ ```bash
492
+ # Build the optimizer
493
+ cargo build -p edgebot-core --bin edgebot-optimize --release
494
+
495
+ # Optimize a model with int8 quantization
496
+ ./target/release/edgebot-optimize \
497
+ --input model.onnx \
498
+ --output model.ebmodel \
499
+ --quantize int8 \
500
+ --fuse-layers
501
+
502
+ # With pruning (magnitude-based, 50% threshold)
503
+ ./target/release/edgebot-optimize \
504
+ --input model.onnx \
505
+ --output model.ebmodel \
506
+ --quantize fp16 \
507
+ --prune magnitude \
508
+ --pruning-threshold 0.5 \
509
+ --device cpu
510
+ ```
511
+
512
+ **Output:** `.ebmodel` bundle containing optimized model + metadata (JSON with embedded binary).
513
+
514
+ **CLI Options:**
515
+ - `--quantize`: none/int8/fp16 (default: none)
516
+ - `--prune`: none/magnitude/structured
517
+ - `--pruning-threshold`: fraction of weights to prune (0.0-1.0)
518
+ - `--fuse-layers`: enable layer fusion (Conv+ReLU, etc.)
519
+ - `--device`: target device (cpu/cuda)
520
+
521
+ **Optimization Stats:** The CLI prints size reduction, speedup estimates, and saves detailed stats in the .ebmodel bundle.
522
+
523
+ ### Webots Simulation
524
+
525
+ The `edgebot-sim` crate provides integration with the Webots robotics simulator for testing AI models on virtual robots in a controlled environment. It offers a safe, Python-like Rust API and supports headless (no-GUI) mode for automated testing and CI.
526
+
527
+ #### Features
528
+
529
+ - **Supervisor control**: Launch simulations, spawn robots, and manipulate the scene.
530
+ - **Sensor access**: Read data from cameras, LiDAR, distance sensors, IMU, GPS, etc.
531
+ - **Headless mode**: Run simulations without a display, perfect for servers and CI.
532
+ - **Remote control**: Connect to a running Webots instance or launch a new one directly from Rust.
533
+ - **Zero-copy**: Efficient memory access to sensor data buffers.
534
+
535
+ #### Usage Example
536
+
537
+ ```rust
538
+ use edgebot_sim::webots::{Supervisor, Robot, Device, WebotsError};
539
+
540
+ fn main() -> Result<(), WebotsError> {
541
+ // Launch Webots in headless mode with a world file
542
+ let supervisor = Supervisor::launch("worlds/warehouse.wbt", true)?;
543
+
544
+ // Spawn a robot from a prototype
545
+ let robot = supervisor.spawn_robot("prototypes/turtlebot3.proto", "test_bot")?;
546
+
547
+ // Step simulation to allow robot to initialize
548
+ supervisor.step(32)?;
549
+
550
+ // Get devices
551
+ let camera = robot.get_device("camera")?.as_camera()?;
552
+ let lidar = robot.get_device("lidar")?.as_lidar()?;
553
+ let wheel_motor = robot.get_device("wheel_left")?;
554
+
555
+ // Enable sensors with appropriate sampling period
556
+ camera.enable(32);
557
+ lidar.enable(32);
558
+
559
+ // Main simulation loop
560
+ for _ in 0..1000 {
561
+ supervisor.step(32)?;
562
+
563
+ // Retrieve camera image
564
+ let image = camera.get_image()?; // RGBA buffer
565
+ // Run inference with edgebot-core here...
566
+
567
+ // Retrieve lidar scan
568
+ let ranges = lidar.get_range_image()?; // Vec<f32>
569
+
570
+ // Simple obstacle avoidance example
571
+ if ranges.iter().any(|&r| r < 0.3) {
572
+ // Stop or reverse
573
+ }
574
+ }
575
+
576
+ // Clean shutdown
577
+ supervisor.terminate()?;
578
+ Ok(())
579
+ }
580
+ ```
581
+
582
+ #### Headless Mode
583
+
584
+ Headless mode runs Webots without a graphical interface, ideal for automated testing and CI pipelines. Set the `WEBOTS_HOME` environment variable to your Webots installation directory:
585
+
586
+ ```bash
587
+ export WEBOTS_HOME=/usr/local/webots
588
+ cargo run --bin my_simulation_test
589
+ ```
590
+
591
+ The `Supervisor::launch` function automatically starts Webots with `--batch` and `--no-rendering` flags.
592
+
593
+ #### Remote Control
594
+
595
+ You can also connect to an already running Webots instance (with remote control enabled on port 1234):
596
+
597
+ ```rust
598
+ let supervisor = Supervisor::connect("localhost", 1234)?;
599
+ ```
600
+
601
+ #### API Reference
602
+
603
+ Key types:
604
+ - `Supervisor`: Main simulation controller. Provides `spawn_robot`, `step`, `get_root`, `load_world`, etc.
605
+ - `Robot`: Handle to a robot in the scene. Provides `get_device`, `get_node`, etc.
606
+ - `Device`: Base device; can be cast to specific device types (`as_camera`, `as_lidar`, etc.).
607
+ - `Node`: Scene tree node for robot and object manipulation.
608
+ - `Field`: Access to node fields (position, rotation, children, etc.).
609
+
610
+ Common methods:
611
+ - `Supervisor::launch(world_path, headless)` -> Launch Webots and connect.
612
+ - `Supervisor::spawn_robot(proto_url, name)` -> Create a new robot from prototype.
613
+ - `Supervisor::step(ms)` -> Advance simulation.
614
+ - `Device::enable(sampling_period)` -> Start sampling a sensor.
615
+ - `Camera::get_image()` -> Get RGBA image bytes.
616
+ - `Lidar::get_range_image()` -> Get distance measurements.
617
+
618
+ For a full list of methods, see the API documentation or the source code in `edgebot-sim/src/webots.rs`.
619
+
620
+ ### ROS2 Bridge
621
+
622
+ The `edgebot-ros2` crate provides ROS2 integration using the `rclrs` crate and standard ROS2 message definitions (`sensor_msgs`, `std_msgs`, `vision_msgs`). It enables publishing and subscribing to ROS2 topics with zero-copy message passing for sensor data (camera images, LiDAR). A YOLO inference example node is included, demonstrating how to subscribe to camera images, run inference with edgebot-core, and publish detection results.
623
+
624
+ #### Running the YOLO Example
625
+
626
+ ```bash
627
+ # Build the example node
628
+ cargo build -p edgebot-ros2 --bin yolo_node --release
629
+
630
+ # Run (requires a ROS2 environment and a camera topic publishing images)
631
+ cargo run -p edgebot-ros2 --bin yolo_node -- --ros-args -p camera_topic:=/camera/color/image_raw
632
+ ```
633
+
634
+ Note: The example currently uses a hardcoded model path `models/yolov8.onnx`. Place a suitable ONNX model in that location or modify the source code to point to your model.
635
+
636
+ ### WebAssembly Runtime
637
+
638
+ The `edgebot-wasm` crate enables deployment of EdgeBot AI models in browser and IoT environments via WebAssembly. It provides:
639
+
640
+ - **Browser support** (`wasm32-unknown-unknown`) with `wasm-bindgen` for JavaScript interop
641
+ - **WASI support** (`wasm32-wasi`) for headless IoT devices
642
+ - Zero-copy memory interfaces for efficient data passing between JS/Rust and WASM
643
+ - Unified API for both targets with runtime selection
644
+
645
+ #### Usage in JavaScript (Browser)
646
+
647
+ ```javascript
648
+ // Import the WASM module (after building with `build-wasm.sh`)
649
+ import init, { JsWasmRuntime } from './edgebot-wasm-browser.js';
650
+
651
+ // Initialize the runtime
652
+ await init();
653
+
654
+ // Create runtime
655
+ const runtime = new JsWasmRuntime();
656
+
657
+ // Load a model (Uint8Array of .ebmodel or .onnx bytes)
658
+ const modelBytes = await fetch('model.ebmodel').then(r => r.arrayBuffer());
659
+ runtime.load_model('yolo', new Uint8Array(modelBytes));
660
+
661
+ // Run inference
662
+ const inputs = [{
663
+ name: 'input',
664
+ data: [/* float32 array */],
665
+ shape: [1, 3, 640, 640]
666
+ }];
667
+ const outputs = runtime.infer('yolo', inputs);
668
+
669
+ console.log('Inference output:', outputs[0].data);
670
+ ```
671
+
672
+ #### Usage in Rust (WASI)
673
+
674
+ ```rust
675
+ use edgebot_wasm::{WasmRuntime, WasmTarget, WasmInferenceInput};
676
+
677
+ fn main() -> Result<(), Box<dyn std::error::Error>> {
678
+ // Create WASI runtime
679
+ let mut runtime = WasmRuntime::new(WasmTarget::Wasi);
680
+
681
+ // Load model from file (WASI filesystem access)
682
+ let model_bytes = std::fs::read("/models/yolo.ebmodel")?;
683
+ runtime.load_model("yolo", model_bytes, None)?;
684
+
685
+ // Prepare input
686
+ let input = WasmInferenceInput {
687
+ name: "input".to_string(),
688
+ data: vec![0.0; 3 * 640 * 640],
689
+ shape: vec![1, 3, 640, 640],
690
+ };
691
+
692
+ // Run inference
693
+ let outputs = runtime.infer("yolo", &[input], None)?;
694
+ println!("Output shape: {:?}", outputs[0].shape);
695
+
696
+ Ok(())
697
+ }
698
+ ```
699
+
700
+ #### Building WASM Modules
701
+
702
+ The `edgebot-wasm/build-wasm.sh` script builds optimized WASM binaries for both targets:
703
+
704
+ ```bash
705
+ # Build browser target (default)
706
+ ./edgebot-wasm/build-wasm.sh browser
707
+
708
+ # Build WASI target
709
+ ./edgebot-wasm/build-wasm.sh wasi
710
+
711
+ # Build both targets
712
+ ./edgebot-wasm/build-wasm.sh all
713
+
714
+ # Debug build
715
+ ./edgebot-wasm/build-wasm.sh browser --debug
716
+
717
+ # With additional size optimizations
718
+ ./edgebot-wasm/build-wasm.sh browser --optimize
719
+ ```
720
+
721
+ Output files are placed in `target/wasm/`:
722
+ - `edgebot-wasm-browser.wasm` - Browser module (requires JS glue code)
723
+ - `edgebot-wasm-wasi.wasm` - WASI standalone module
724
+
725
+ #### Requirements
726
+
727
+ Add to your Cargo.toml:
728
+ ```toml
729
+ [dependencies]
730
+ edgebot-wasm = { path = "edgebot-wasm" }
731
+
732
+ # For browser builds
733
+ [target.'cfg(target_arch = "wasm32")'.dependencies]
734
+ wasm-bindgen = "0.2"
735
+ ```
736
+
737
+ Install wasm32 targets:
738
+ ```bash
739
+ rustup target add wasm32-unknown-unknown wasm32-wasi
740
+ ```
741
+
742
+ #### API Reference
743
+
744
+ **Core Types:**
745
+ - `WasmRuntime`: Unified runtime for model loading and inference
746
+ - `WasmTarget`: Platform target (`Browser` or `Wasi`)
747
+ - `WasmInferenceInput`: Input tensor with name, data (Vec<f32>), and shape
748
+ - `WasmInferenceOutput`: Inference result with name, data, and shape
749
+
750
+ **Key Methods:**
751
+ - `WasmRuntime::new(target)` - Create runtime for specific platform
752
+ - `load_model(name, bytes)` - Load model from bytes (requires .ebmodel or supported format)
753
+ - `infer(model_name, inputs)` - Run inference with vector of inputs
754
+ - `list_models()` - List loaded model names
755
+ - `unload_model(name)` - Free model resources
756
+
757
+ **Browser-Specific:**
758
+ - `JsWasmRuntime`: Web-friendly runtime (automatically selected via `#[cfg(target_arch = "wasm32")]`)
759
+ - `new()` constructor available from JavaScript
760
+ - Methods return `Result<..., JsValue>` for proper error handling in JS
761
+
762
+ **WASI-Specific:**
763
+ - `WasiJsRuntime::new()` - Create WASI runtime
764
+ - `load_model_from_path(name, path)` - Load model from filesystem
765
+ - Automatic support for WASI environment (files, stdin/stdout)
766
+
767
+ #### Performance Notes
768
+
769
+ - Browser target uses WGPU for GPU acceleration (via Burn's wgpu backend)
770
+ - WASI target uses CPU-optimized backends (Autocast, Tch)
771
+ - Zero-copy memory interfaces minimize data marshaling overhead
772
+ - Release builds with `opt-level = "z"` produce ~50% smaller WASM binaries
773
+ - LTO and codegen-units=1 further reduce size
774
+
775
+ ## License
776
+
777
+ MIT OR Apache-2.0. See [LICENSE](LICENSE) for details.
778
+
779
+ ## Contributing
780
+
781
+ We welcome contributions! See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.
782
+
783
+ ## Documentation
784
+
785
+ Comprehensive documentation is available as a book:
786
+
787
+ - **Online**: [https://edgebot-ai.github.io/edgebot-ai/](https://edgebot-ai.github.io/edgebot-ai/)
788
+ - **Local**: `docs/book/` (markdown source) or `cargo doc --workspace --open`
789
+
790
+ The documentation covers:
791
+
792
+ - Quickstart guide
793
+ - ROS2 integration
794
+ - WebAssembly deployment
795
+ - Pro workflow and licensing
796
+ - API reference
797
+
798
+ ### Building the Docs
799
+
800
+ ```bash
801
+ # Install mdbook (for book build)
802
+ cargo install mdbook
803
+
804
+ # Build the book
805
+ mdbook build docs/book
806
+
807
+ # Serve locally
808
+ mdbook serve docs/book --open
809
+ ```
810
+
811
+ The book is automatically published to GitHub Pages on pushes to `main`.
812
+
813
+ ---
814
+
815
+ **Status:** Initial release candidate. API stabilizing.
docs/readmes/esp-robot-README.md ADDED
@@ -0,0 +1,281 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # esp-robot
2
+ **Repository:** Julien-ser/esp-robot
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** C++
7
+ **Last Updated:** 2026-03-13T02:29:36Z
8
+
9
+ ---
10
+
11
+ # ESP-32 Robot with Obstacle Avoidance
12
+
13
+ An autonomous robot platform based on ESP32 with dual motor control and ultrasonic obstacle detection.
14
+
15
+ ## Hardware Components
16
+
17
+ - ESP32 (esp32dev board)
18
+ - Dual motor driver (L298N, TB6612, or compatible)
19
+ - 2x DC motors with wheels
20
+ - HC-SR04 ultrasonic sensor
21
+ - Robot chassis kit
22
+ - Power supply (7-12V for motors, 5V for ESP32)
23
+
24
+ ## GPIO Pin Configuration
25
+
26
+ See `include/pin_config.h` for complete pin definitions.
27
+
28
+ | Component | Signal | ESP32 Pin | Notes |
29
+ |--------------|---------|-----------|------------------------------|
30
+ | Motor A | ENA | GPIO13 | PWM speed control (ch0) |
31
+ | Motor A | IN1 | GPIO12 | Direction |
32
+ | Motor A | IN2 | GPIO14 | Direction |
33
+ | Motor B | ENB | GPIO27 | PWM speed control (ch1) |
34
+ | Motor B | IN3 | GPIO26 | Direction |
35
+ | Motor B | IN4 | GPIO25 | Direction |
36
+ | Ultrasonic | TRIG | GPIO23 | 5V tolerant |
37
+ | Ultrasonic | ECHO | GPIO22 | Requires voltage divider |
38
+
39
+ **Important:** ECHO pin needs a voltage divider (2.2kΩ + 3.3kΩ) to convert 5V→3.3V.
40
+
41
+ ## Project Structure
42
+
43
+ ```
44
+ esp-robot/
45
+ ├── include/ # Header files
46
+ │ └── pin_config.h # Pin definitions and hardware constants
47
+ ├── src/ # Source code
48
+ │ ├── main.cpp
49
+ │ ├── Motor.cpp/h # Individual motor control via PWM
50
+ │ ├── MotorDriver.cpp/h # Differential drive control
51
+ │ ├── UltrasonicSensor.cpp/h # HC-SR04 interface
52
+ │ ├── ObstacleDetector.cpp/h # Filtered distance reading
53
+ │ ├── SensorArray.cpp/h # Multi-sensor management
54
+ │ ├── calibration.cpp/h # Motor speed calibration
55
+ │ └── Robot.cpp/h # Main state machine (planned)
56
+ ├── test/ # Unit tests
57
+ │ ├── motor_test.cpp
58
+ │ └── sensor_test.cpp
59
+ ├── platformio.ini # PlatformIO configuration
60
+ ├── TASKS.md # Development progress
61
+ └── README.md # This file
62
+ ```
63
+
64
+ ## Calibration Procedure
65
+
66
+ The motors may have different characteristics, so calibration is essential for accurate driving and turning.
67
+
68
+ ### What Calibration Does
69
+
70
+ The calibration routine:
71
+ 1. Tests the robot at multiple PWM values (50-255)
72
+ 2. Measures actual speed (cm/s) for each motor
73
+ 3. Stores a PWM-to-speed lookup table in NVS (non-volatile storage)
74
+ 4. Enables speed compensation: you can request a speed in cm/s and the system converts to the correct PWM
75
+
76
+ ### How to Run Calibration
77
+
78
+ 1. Place the robot on a clear, level surface with a measured distance (default: 50cm)
79
+ 2. Upload the code with calibration support
80
+ 3. Open serial monitor at 115200 baud
81
+ 4. Send the `cal` command (or trigger calibration mode via code)
82
+ 5. Follow the prompts:
83
+ - For each PWM test value, place robot at start marker
84
+ - Press any key to start the test run
85
+ - Robot will drive forward at the test PWM
86
+ - Press any key when it reaches the end marker
87
+ - The system records time and calculates speed
88
+ 6. Repeat for both motors (left then right)
89
+ 7. Data is automatically saved to NVS and persists across reboots
90
+
91
+ ### Calibration Data Storage
92
+
93
+ Data is stored in ESP32 NVS under namespace `robot_cal` with key `cal_data`. To clear calibration:
94
+ - Send `cal_clear` command, OR
95
+ - Call `calibration.clear()` in code
96
+
97
+ ### Using Calibrated Speeds
98
+
99
+ After calibration, use:
100
+ - `MotorDriver::driveForwardCalibrated(float speed_cm_s)` - drive at exact cm/s
101
+ - `MotorDriver::turnCalibrated(uint8_t baseSpeed, int16_t degrees)` - precise turns using speed-based differential
102
+ - `calibration.getPWMForSpeed('L' or 'R', desired_speed)` - manual PWM lookup
103
+
104
+ ## Current Progress
105
+
106
+ ✅ **Phase 1 Complete:**
107
+ - [x] GPIO pin assignments defined in `include/pin_config.h`
108
+ - [x] PlatformIO project initialization
109
+ - [x] Directory structure and skeleton files
110
+ - [x] Serial monitor configuration and blink test
111
+
112
+ ✅ **Phase 2 Complete:**
113
+ - [x] Motor class with PWM speed control
114
+ - [x] MotorDriver class with movement methods
115
+ - [x] Unit tests for motor control
116
+ - [x] Calibration routine with NVS storage
117
+
118
+ ✅ **Phase 3 Complete:**
119
+ - [x] UltrasonicSensor class (NewPing library)
120
+ - [x] ObstacleDetector with filtering and hysteresis
121
+ - [x] Multi-sensor support via SensorArray
122
+ - [x] Sensor tests (unit tests with mock ISonar interface)
123
+
124
+ ✅ **Phase 4 Complete:**
125
+ - [x] Robot state machine (IDLE, DRIVING, AVOIDING, STOPPED)
126
+ - [x] Obstacle avoidance algorithm
127
+ - [x] Serial remote control with safety interlock
128
+ - [x] Debug dashboard (1Hz serial output)
129
+ - [x] Auto-start after 5s delay
130
+
131
+ ✅ **Phase 5 Complete:**
132
+ - [x] Final integration test suite (`test/robot_integration_test.cpp`) with test doubles
133
+ - [x] Full system documentation and assembly guide
134
+
135
+ ---
136
+
137
+ ### Testing
138
+
139
+ The project includes a comprehensive integration test suite that validates the robot's state machine, motor control, and obstacle avoidance logic using test doubles.
140
+
141
+ **Run tests with PlatformIO:**
142
+
143
+ ```bash
144
+ pio test
145
+ ```
146
+
147
+ The integration test uses mock implementations of `IMotorDriver` and `IObstacleDetector` to simulate sensor readings and verify motor outputs. It covers:
148
+ - Initialization and IDLE state
149
+ - Starting autonomous driving
150
+ - Obstacle detection triggering avoidance
151
+ - Full avoidance sequence (stop, reverse, pivot turn, resume)
152
+ - Emergency stop and reset functionality
153
+ - Manual command rejection when not in IDLE
154
+
155
+ Test results are output to the serial monitor at 115200 baud when running on actual hardware, or displayed in the console for native builds.
156
+
157
+ ---
158
+
159
+ ## Serial Remote Control
160
+
161
+ The robot can be controlled via serial commands (115200 baud). Commands are single characters:
162
+
163
+ | Command | Description | Requirements |
164
+ |---------|-------------|--------------|
165
+ | w | Drive forward at default speed | Robot in IDLE state |
166
+ | s | Drive backward | Robot in IDLE state |
167
+ | a | Turn left (differential, ~90°) | Robot in IDLE state |
168
+ | d | Turn right (differential, ~90°) | Robot in IDLE state |
169
+ | q | Pivot counter-clockwise | Robot in IDLE state |
170
+ | e | Pivot clockwise | Robot in IDLE state |
171
+ | x | Emergency stop (brake) | Always accepted |
172
+ | r | Reset from STOPPED to IDLE | Only from STOPPED |
173
+ | ? or h | Show this help | Always accepted |
174
+
175
+ **Safety Interlock:** Movement commands (w, s, a, d, q, e) are only accepted when the robot is in IDLE state. This prevents accidental commands during autonomous operation.
176
+
177
+ **Auto-Start:** If the robot remains IDLE for 5 seconds, it automatically enters autonomous driving mode. Send commands before then to retain manual control.
178
+
179
+ ### Example Usage
180
+ ```bash
181
+ # Open serial monitor
182
+ screen /dev/ttyUSB0 115200
183
+
184
+ # Press 'w' to drive forward (must be in IDLE)
185
+ # Press 'x' for emergency stop
186
+ # Press 'r' to reset after emergency stop
187
+ ```
188
+
189
+ ## Debug Dashboard
190
+
191
+ The system outputs status information every 1 second:
192
+ - Current state (IDLE, DRIVING, AVOIDING, STOPPED)
193
+ - Movement status
194
+ - Distance reading (mm)
195
+ - Obstacle detection status
196
+
197
+ Motor speeds can be added by connecting an ADC pin for battery monitoring (future enhancement).
198
+
199
+ ## Setup Instructions
200
+
201
+ 1. Install PlatformIO (VS Code extension or CLI)
202
+ 2. Create `platformio.ini` configured for `esp32dev` board
203
+ 3. Install required libraries: `NewPing`, `PID`, `AsyncTCP` (optional)
204
+ 4. Connect hardware according to pin configuration
205
+ 5. Build and flash to ESP32
206
+
207
+ ## Motor Control
208
+
209
+ - PWM frequency: 5kHz
210
+ - Speed range: 0-255 (8-bit)
211
+ - Direction control via two GPIO pins per motor
212
+ - Functionality: forward, backward, stop, brake, variable speed
213
+
214
+ ## Ultrasonic Sensor
215
+
216
+ - Model: HC-SR04
217
+ - Max range: 500cm (5m)
218
+ - Trigger pulse: 10µs
219
+ - Timeout: 30ms
220
+ - Pin TRIG: output (5V compatible)
221
+ - Pin ECHO: input with 5V→3.3V voltage divider
222
+
223
+ ## Obstacle Avoidance
224
+
225
+ The robot implements a simple state machine with avoidance behavior:
226
+ 1. Drive forward continuously
227
+ 2. Monitor front distance
228
+ 3. If obstacle within 50cm: stop, reverse 0.5s, pivot turn 90°
229
+ 4. Resume driving
230
+
231
+ ## Calibration Procedure
232
+
233
+ The motors may have different characteristics, so calibration is essential for accurate driving and turning.
234
+
235
+ ### What Calibration Does
236
+
237
+ The calibration routine:
238
+ 1. Tests the robot at multiple PWM values (50-255)
239
+ 2. Measures actual speed (cm/s) for each motor
240
+ 3. Stores a PWM-to-speed lookup table in NVS (non-volatile storage)
241
+ 4. Enables speed compensation: you can request a speed in cm/s and the system converts to the correct PWM
242
+
243
+ ### How to Run Calibration
244
+
245
+ 1. Place the robot on a clear, level surface with a measured distance (default: 50cm)
246
+ 2. Upload the code with calibration support
247
+ 3. Open serial monitor at 115200 baud
248
+ 4. Send the `cal` command (or trigger calibration mode via code)
249
+ 5. Follow the prompts:
250
+ - For each PWM test value, place robot at start marker
251
+ - Press any key to start the test run
252
+ - Robot will drive forward at the test PWM
253
+ - Press any key when it reaches the end marker
254
+ - The system records time and calculates speed
255
+ 6. Repeat for both motors (left then right)
256
+ 7. Data is automatically saved to NVS and persists across reboots
257
+
258
+ ### Calibration Data Storage
259
+
260
+ Data is stored in ESP32 NVS under namespace `robot_cal` with key `cal_data`. To clear calibration:
261
+ - Send `cal_clear` command, OR
262
+ - Call `calibration.clear()` in code
263
+
264
+ ### Using Calibrated Speeds
265
+
266
+ After calibration, use:
267
+ - `MotorDriver::driveForwardCalibrated(float speed_cm_s)` - drive at exact cm/s
268
+ - `MotorDriver::turnCalibrated(uint8_t baseSpeed, int16_t degrees)` - precise turns using speed-based differential
269
+ - `calibration.getPWMForSpeed('L' or 'R', desired_speed)` - manual PWM lookup
270
+
271
+ ## Power Requirements
272
+
273
+ - ESP32: 5V via USB or regulator (draws ~200-500mA)
274
+ - Motors: 7-12V depending on motor specs (up to 1A each under load)
275
+ - Motor driver: Must match motor voltage, handle peak current
276
+
277
+ Recommended: Separate power for motors and logic to avoid noise.
278
+
279
+ ## License
280
+
281
+ TBD
docs/readmes/flash-quote-api-README.md ADDED
@@ -0,0 +1,142 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # flash-quote-api
2
+ **Repository:** Julien-ser/flash-quote-api
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-27T23:01:28Z
8
+
9
+ ---
10
+
11
+ # flash-quote-api
12
+
13
+ A tiny FastAPI service that returns random inspirational quotes.
14
+
15
+ ## Architecture
16
+
17
+ **Storage Design:**
18
+ - Quotes are stored in `quotes_data.json` (JSON array with id, text, author, category)
19
+ - On startup, the entire JSON file is loaded into memory as a Python list
20
+ - All API endpoints read from this in-memory array for maximum performance
21
+ - This approach provides O(1) random access and O(n) category filtering with minimal overhead
22
+
23
+ **Trade-offs:**
24
+ - ✅ Fast read performance (no database queries)
25
+ - ✅ Simple implementation and maintenance
26
+ - ✅ No external dependencies beyond FastAPI
27
+ - ⚠️ Data is read-only at runtime (modify JSON file to update quotes)
28
+ - ⚠️ Memory usage scales with quote count (negligible for <1000 quotes)
29
+
30
+ ## Quick Start
31
+
32
+ ### Local Development
33
+
34
+ ```bash
35
+ # Install dependencies
36
+ pip install fastapi uvicorn
37
+
38
+ # Run the server
39
+ uvicorn main:app --reload
40
+ ```
41
+
42
+ ### Docker Deployment
43
+
44
+ ```bash
45
+ # Build the Docker image
46
+ docker build -t flash-quote-api .
47
+
48
+ # Run the container
49
+ docker run -p 8000:8000 flash-quote-api
50
+ ```
51
+
52
+ The API will be available at http://localhost:8000
53
+
54
+ ## API Documentation
55
+
56
+ The API provides three endpoints for accessing inspirational quotes. FastAPI automatically generates interactive OpenAPI/Swagger documentation at:
57
+
58
+ - **Swagger UI**: http://localhost:8000/docs
59
+ - **ReDoc**: http://localhost:8000/redoc
60
+
61
+ ### Endpoints
62
+
63
+ #### GET / (Root)
64
+
65
+ Returns a single random quote.
66
+
67
+ **Response:**
68
+ ```json
69
+ {
70
+ "id": 42,
71
+ "text": "The only way to do great work is to love what you do.",
72
+ "author": "Steve Jobs",
73
+ "category": "motivation"
74
+ }
75
+ ```
76
+
77
+ #### GET /quotes
78
+
79
+ Returns all quotes, with optional category filtering.
80
+
81
+ **Query Parameters:**
82
+ - `category` (optional): Filter quotes by category (e.g., "motivation", "life", "inspiration")
83
+
84
+ **Example (all quotes):**
85
+ ```json
86
+ [
87
+ {
88
+ "id": 1,
89
+ "text": "The only way to do great work is to love what you do.",
90
+ "author": "Steve Jobs",
91
+ "category": "motivation"
92
+ },
93
+ {
94
+ "id": 2,
95
+ "text": "Life is what happens when you're busy making other plans.",
96
+ "author": "John Lennon",
97
+ "category": "life"
98
+ }
99
+ ]
100
+ ```
101
+
102
+ **Example (filtered by category=motivation):**
103
+ ```json
104
+ [
105
+ {
106
+ "id": 1,
107
+ "text": "The only way to do great work is to love what you do.",
108
+ "author": "Steve Jobs",
109
+ "category": "motivation"
110
+ }
111
+ ]
112
+ ```
113
+
114
+ #### GET /quote/{quote_id}
115
+
116
+ Returns a specific quote by its unique ID.
117
+
118
+ **Path Parameters:**
119
+ - `quote_id`: The unique numeric identifier of the quote
120
+
121
+ **Example response:**
122
+ ```json
123
+ {
124
+ "id": 42,
125
+ "text": "The only way to do great work is to love what you do.",
126
+ "author": "Steve Jobs",
127
+ "category": "motivation"
128
+ }
129
+ ```
130
+
131
+ **Error responses:**
132
+ - `404 Not Found`: When the specified quote ID does not exist
133
+
134
+ ## Project Status
135
+
136
+ **Phase 1: Planning & Setup**
137
+ - ✅ Quote data structure defined (JSON with id, text, author, category)
138
+ - ✅ Storage method: JSON file loaded into memory at startup
139
+ - ✅ Initialize FastAPI project with uv/pip
140
+ - ✅ Set up tests directory and add unit tests
141
+
142
+ See [TASKS.md](TASKS.md) for full development roadmap.
docs/readmes/football-management-sim-README.md ADDED
@@ -0,0 +1,577 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # football-management-sim
2
+ **Repository:** Julien-ser/football-management-sim
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** TypeScript
7
+ **Last Updated:** 2026-03-17T08:01:48Z
8
+
9
+ ---
10
+
11
+ # Football Manager Simulator
12
+
13
+ **Version 0.1.0 (Beta)** | [Release Notes](docs/RELEASE_NOTES.md) | [User Manual](docs/USER_MANUAL.md)
14
+
15
+ A strategic football management simulation game where you take charge of a football club and lead it to domestic and European glory.
16
+
17
+ ## 🎯 Mission
18
+
19
+ Build, manage, and guide your football club through multiple seasons, balancing squad development, financial management, tactical innovation, and competition success across domestic leagues, cups, and European tournaments.
20
+
21
+ ## 🏆 Beta Release Ready (v0.1.0)
22
+
23
+ The game is now **feature-complete** and ready for beta testing. All core systems implemented, tested, and optimized.
24
+
25
+ ### ✅ Current Status
26
+
27
+ **Phase 1-3: All Complete** ✅
28
+
29
+ All major features delivered:
30
+
31
+ - Match simulation engine with AI team behavior
32
+ - Comprehensive tactics & formations system
33
+ - Full transfer market with scouting & negotiation
34
+ - Multi-competition support (domestic + European)
35
+ - Complete React-based UI with HUD, match day, menus
36
+ - Audio system with crowds, events, and UI sounds
37
+ - Save/load with 10 slots and auto-save
38
+
39
+ **Phase 4: Testing & Deployment - In Progress**
40
+
41
+ - ✅ Automated test suite: **348 tests** with **89.59%** coverage
42
+ - ✅ Beta testing infrastructure and performance benchmarks
43
+ - ✅ Gameplay balancing complete
44
+ - ✅ **Release build ready** (`npm run build:web` → `dist-web/`)
45
+
46
+ **Phase 4.4: Release Builds & Documentation** ⬜ **In Progress**
47
+
48
+ - ✅ Web production build created
49
+ - ✅ User manual written
50
+ - ✅ Release notes prepared
51
+ - ⬜ Promotional assets (screenshots, icons)
52
+ - ⬜ Final deployment checklist
53
+
54
+ See [Release Notes](docs/RELEASE_NOTES.md) for full details and [User Manual](docs/USER_MANUAL.md) for gameplay guide.
55
+
56
+ ---
57
+
58
+ ## 🚀 Quick Start
59
+
60
+ ### Play Now (No Installation)
61
+
62
+ 1. Download or clone this repository
63
+ 2. Navigate to `dist-web/` folder
64
+ 3. Open `index.html` in any modern browser
65
+ - OR serve with: `npx serve dist-web` or `python -m http.server -d dist-web`
66
+
67
+ That's it! No build tools or dependencies needed to play.
68
+
69
+ ### For Development & Testing
70
+
71
+ ```bash
72
+ # Install dependencies
73
+ npm install
74
+
75
+ # Run development server (hot reload)
76
+ npm run dev
77
+
78
+ # Build for production
79
+ npm run build:web # Build React frontend → dist-web/
80
+ npm run build # Build Node backend → dist/
81
+
82
+ # Run tests
83
+ npm test # Unit tests (Jest)
84
+ npm run test:e2e # E2E tests (Cypress)
85
+ npm run test:all # All tests
86
+ npm run benchmark # Performance benchmarks
87
+
88
+ # Preview production build locally
89
+ npm run preview
90
+ ```
91
+
92
+ **Requirements:** Node.js 18+, npm, modern browser
93
+
94
+ ---
95
+
96
+ ## 🎮 Core Features
97
+
98
+ ### Match Simulation Engine
99
+
100
+ - Full 90+ minute simulation in **<5 seconds**
101
+ - Event-driven mechanics: goals, cards, injuries, penalties, corners, fouls
102
+ - Team AI with tactical awareness
103
+ - Real-time stats: possession, shots, passes, fouls, corners, offsides
104
+ - Performance: **27.2ms per match minute** (4x faster than 100ms target)
105
+
106
+ ### Tactics & Formation System
107
+
108
+ - 6 preset formations (4-4-2, 4-3-3, 3-5-2, etc.)
109
+ - Customizable tactical dimensions:
110
+ - Mentality (Defensive ↔ Attacking)
111
+ - Pressing (Low ↔ High)
112
+ - Passing Style (Short ↔ Long)
113
+ - Width (Narrow ↔ Wide)
114
+ - Defensive Line (Low ↔ High)
115
+ - In-match tactical changes
116
+ - Save/load custom tactic presets
117
+
118
+ ### Transfer Market
119
+
120
+ - **Player Search:** Filter by position, rating, age, nationality, salary
121
+ - **Scouting System:** Detailed reports with ratings, potential, strengths/weaknesses
122
+ - **Negotiation:** Bidding, contract offers (wages, bonuses, length), counter-offers
123
+ - **AI Clubs:** Computer teams actively participate in market
124
+ - **Squad Registration:** Competition squad management with validation
125
+ - **Budget Management:** Track fees, wages, ensure financial health
126
+
127
+ ### Competitions & Calendar
128
+
129
+ - **Domestic League:** Round-robin format with promotion/relegation
130
+ - **Domestic Cups:** Single-elimination tournaments
131
+ - **European Competitions:**
132
+ - UEFA Champions League
133
+ - UEFA Europa League
134
+ - UEFA Conference League
135
+ - **Calendar System:** Fixture generation, qualification rules, progression
136
+ - **Match Scheduling:** Realistic season calendar
137
+
138
+ ### Interface & Experience
139
+
140
+ - **Main Menu:** Start career, load game, settings
141
+ - **Club Selection:** Browse and choose from playable clubs
142
+ - **Game HUD:** Quick access to all management functions
143
+ - **Match Day UI:** Live commentary, tactical overlay, animated stats
144
+ - **Settings:** Graphics quality, audio controls, match speed, auto-save
145
+ - **Save/Load:** 10 slots + auto-save with configurable intervals
146
+
147
+ ### Audio System
148
+
149
+ - Crowd ambience and match event sounds
150
+ - UI feedback sounds
151
+ - Configurable volume controls (master, music, effects)
152
+ - Mute support
153
+
154
+ ---
155
+
156
+ ## 🏗️ Architecture
157
+
158
+ **Tech Stack:**
159
+
160
+ | Component | Technology |
161
+ | ------------- | ----------------------- |
162
+ | Frontend UI | React 18 + TypeScript 5 |
163
+ | Build Tool | Vite 5 |
164
+ | State/Events | RxJS 7 |
165
+ | Backend (CLI) | Node.js + TypeScript |
166
+ | Database | SQLite (better-sqlite3) |
167
+ | Testing | Jest + Cypress |
168
+ | Linting | ESLint + Prettier |
169
+
170
+ **Architecture:** Layered clean architecture
171
+
172
+ ```
173
+ Presentation Layer (React UI)
174
+
175
+ Application Logic (Match Engine, Tactics, Transfers, etc.)
176
+
177
+ Data Layer (SQLite + Models)
178
+ ```
179
+
180
+ **Future Roadmap:**
181
+
182
+ - Youth academy and training systems
183
+ - Staff management
184
+ - Club infrastructure upgrades
185
+ - Historical statistics
186
+ - Multiplayer cloud saves
187
+ - Mobile responsiveness
188
+
189
+ See [Technology ADR](docs/ADR-001-technology-stack.md) for detailed rationale.
190
+
191
+ ---
192
+
193
+ ## 📦 Build & Distribution
194
+
195
+ ### Production Builds
196
+
197
+ **Web Version (Current):**
198
+
199
+ ```bash
200
+ npm run build:web # Creates dist-web/ with optimized bundles
201
+ ```
202
+
203
+ Build outputs:
204
+
205
+ - `dist-web/index.html` - Entry point (30 KB HTML)
206
+ - `dist-web/assets/main-*.css` - Styles (30.59 KB, 5.59 KB gzipped)
207
+ - `dist-web/assets/main-*.js` - JavaScript (230.62 KB, 69.13 KB gzipped)
208
+
209
+ **Total Initial Load:** ~300 KB (very fast even on mobile)
210
+
211
+ **Deployment:** Upload entire `dist-web/` folder to any static hosting service:
212
+
213
+ - GitHub Pages, Netlify, Vercel (drag & drop)
214
+ - AWS S3 + CloudFront
215
+ - Any web server with static file support
216
+
217
+ **Desktop Packages:** Not yet implemented. Would require Electron packaging.
218
+
219
+ ---
220
+
221
+ ## 🧪 Testing & Quality
222
+
223
+ ### Automated Test Suite
224
+
225
+ - **348 total tests** covering all core systems
226
+ - **Coverage:**
227
+ - Statements: 89.59%
228
+ - Branches: 81.34%
229
+ - Functions: 87.11%
230
+ - Lines: 91.01%
231
+
232
+ **Test Types:**
233
+
234
+ - Unit tests for models, utilities, game logic
235
+ - Integration tests for workflows (match simulation, transfers, calendar)
236
+ - E2E tests with Cypress for UI flows
237
+ - Performance benchmarks (match speed, memory usage)
238
+
239
+ ### Performance Benchmarks
240
+
241
+ | Metric | Target | Achieved |
242
+ | ---------------- | ---------- | ----------------- |
243
+ | Match simulation | <100ms/min | **27.2ms/min** ✅ |
244
+ | Memory usage | <500MB | **142MB** ✅ |
245
+ | Test coverage | ≥80% | **>89%** ✅ |
246
+ | Build size | <500KB | **~300KB** ✅ |
247
+
248
+ ---
249
+
250
+ ## 📚 Documentation
251
+
252
+ All documentation is in the `docs/` folder:
253
+
254
+ - **[Game Design Document (GDD)](docs/GDD.md)** - Complete design spec with features, wireframes, ER diagrams
255
+ - **[User Manual](docs/USER_MANUAL.md)** - Comprehensive gameplay guide for players
256
+ - **[Release Notes](docs/RELEASE_NOTES.md)** - Version history and changes
257
+ - **[Technology ADR](docs/ADR-001-technology-stack.md)** - Architecture decisions and rationale
258
+ - **[Beta Testing Guide](docs/BETA_TESTING.md)** - Testing procedures and templates
259
+ - **[Beta Test Report](BETA_TEST_REPORT.md)** - Results and feedback summary
260
+ - **[ER Diagram](docs/ER-diagram.md)** - Database schema visualization
261
+
262
+ **Inline Documentation:**
263
+
264
+ - All public functions and classes have TypeScript docstrings
265
+ - Code is self-documenting with clear naming conventions
266
+
267
+ ---
268
+
269
+ ## 💾 Save System
270
+
271
+ - 10 manual save slots
272
+ - Auto-save after every match (configurable interval)
273
+ - Save file format: JSON (human-readable)
274
+ - Browser localStorage for persistence (local development)
275
+ - File-based save system for production (via Node.js backend or future implementation)
276
+
277
+ ---
278
+
279
+ ## 🤝 Contributing
280
+
281
+ This is currently a solo project. Contributions not accepted unless specified in future.
282
+
283
+ **However, feedback is welcome:**
284
+
285
+ - Bug reports via GitHub Issues
286
+ - Feature suggestions via GitHub Discussions
287
+ - Beta testing sign-ups (watch for announcements)
288
+
289
+ ---
290
+
291
+ ## 📋 Project Status
292
+
293
+ **Phases 1-3: ✅ Complete**
294
+
295
+ - ✅ Planning, design, and tech stack selection
296
+ - ✅ Core systems: match engine, tactics, transfers, competitions
297
+ - ✅ Complete UI: HUD, match day, menus, settings
298
+ - ✅ Audio, animations, and polish
299
+
300
+ **Phase 4: Testing & Deployment - 75% Complete**
301
+
302
+ - ✅ Automated tests (348 tests)
303
+ - ✅ Beta testing infrastructure
304
+ - ✅ Performance optimization
305
+ - ⬜ Promotional assets (in progress)
306
+ - ⬜ Final deployment checklist
307
+
308
+ **Phase 4.4: Release Builds & Documentation** - Current task
309
+
310
+ Web build is ready, documentation created. Completing promotional assets and final verification.
311
+
312
+ ---
313
+
314
+ ## 🔮 Roadmap
315
+
316
+ **Completed:**
317
+
318
+ 1. ✅ Phase 1: Planning & Setup
319
+ 2. ✅ Phase 2: Core Game Systems
320
+ 3. ✅ Phase 3: UI/UX & Polish
321
+ 4. 🔄 Phase 4: Testing & Deployment (near completion)
322
+
323
+ **Upcoming (Post-Beta):**
324
+
325
+ - Phase 5: Quality of Life improvements
326
+ - Phase 6: Advanced features (youth academy, training, staff)
327
+ - Phase 7: Community features and expansion
328
+
329
+ See [TASKS.md](TASKS.md) for detailed task tracking.
330
+
331
+ ---
332
+
333
+ ## 📄 License
334
+
335
+ To be determined - currently proprietary for beta testing.
336
+
337
+ ---
338
+
339
+ ## 🙏 Acknowledgments
340
+
341
+ - **Testing:** Beta testers who provided invaluable feedback
342
+ - **Tools:** Jest, Cypress, Vite, React, TypeScript communities
343
+ - **Inspiration:** Classic football management games (Football Manager, Championship Manager)
344
+
345
+ ---
346
+
347
+ ## 🔗 Quick Links
348
+
349
+ - **Play:** Open `dist-web/index.html` in browser
350
+ - **Report Bug:** [GitHub Issues](.github/ISSUE_TEMPLATE/bug_report.md)
351
+ - **Read Docs:** [docs/](docs/)
352
+ - **View Roadmap:** [TASKS.md](TASKS.md)
353
+ - **Source Code:** `src/` directory
354
+ - **Tests:** `src/**/*.test.ts`, `cypress/e2e/`
355
+
356
+ ---
357
+
358
+ **Enjoy your managerial career!** ⚽🏆
359
+
360
+ _Lead your club to domestic and European glory!_
361
+
362
+ ## 🎮 Core Features
363
+
364
+ - **Match Simulation Engine:** Full 90+ minute simulation in <5 seconds with realistic event-driven mechanics including:
365
+ - Goals, own goals, penalties, and missed penalties
366
+ - Yellow and red cards (including second yellow → red)
367
+ - Injuries and substitutions with AI decision-making
368
+ - Real-time statistics: possession, shots, passes, fouls, corners, offsides
369
+ - Tactical influence on match outcomes (formation, mentality, pressing, passing style)
370
+ - Event streaming via RxJS for live commentary and UI updates
371
+ - **Tactics System:** Formation editor (4-4-2, 4-3-3, etc.), team/player instructions
372
+ - **Transfer Market System (Phase 2 Complete):** Comprehensive player transfer management including:
373
+ - **Player Search & Filtering:** Search by position, rating, age, nationality, salary, contract expiry
374
+ - **Scouting System:** Hire scouts with region expertise, generate detailed reports with ratings, potential, strengths/weaknesses, and recommendations
375
+ - **Bidding & Negotiation:** Place bids, accept/reject/counter offers, contract negotiation with salary, bonuses, and contract length
376
+ - **AI Club Behavior:** Computer-controlled clubs actively participate in the transfer market, identifying squad needs and pursuing targets
377
+ - **Squad Registration:** Register players for competitions with position minimums, jersey number uniqueness, and captain/vice-captain assignments
378
+ - **Budget Management:** Track transfers fees, wages, and ensure financial viability with buffer requirements
379
+ - **Competitions:** Domestic leagues, cups, and UEFA Champions League/Europa League/Conference League
380
+ - **Squad Management:** Rosters, contracts, player development, injuries
381
+ - **Club Finances:** Budgeting, revenue streams (matchday, TV, commercial), wage management
382
+ - **Youth Academy:** Recruitment, development, promotion pathway
383
+ - **Audio System:** Complete audio integration including crowd ambience, match event sounds (goals, cards, whistles), and UI feedback. Audio settings (master volume, music, sound effects, mute) are configurable in the Settings panel.
384
+ - **Animations:** The interface includes smooth CSS animations for match commentary events, UI transitions, and interactive elements to enhance the visual experience.
385
+
386
+ ## 🏗️ Architecture
387
+
388
+ The project follows a layered architecture:
389
+
390
+ ```
391
+ Presentation Layer (Console UI → Future: React Web UI)
392
+
393
+ Application Logic (Match Engine, Tactics, Transfer Market, etc.)
394
+
395
+ Data Layer (SQLite database, models, serialization)
396
+ ```
397
+
398
+ ### Technology Stack
399
+
400
+ **Core:**
401
+
402
+ - **Language:** TypeScript 5.x
403
+ - **Runtime:** Node.js 18+ (LTS)
404
+ - **Database:** SQLite with `better-sqlite3` driver
405
+ - **State/Events:** RxJS for reactive event handling
406
+
407
+ **Development Tools:**
408
+
409
+ - **Build:** TypeScript Compiler (`tsc`)
410
+ - **Testing:** Jest + ts-jest
411
+ - **Linting:** ESLint + @typescript-eslint
412
+ - **Formatting:** Prettier
413
+ - **CI/CD:** GitHub Actions (multi-version Node.js testing)
414
+
415
+ **Future (Phase 3):**
416
+
417
+ - **UI Framework:** React 18+ (for graphical interface)
418
+ - **Styling:** Tailwind CSS or CSS Modules
419
+ - **Routing:** React Router
420
+ - **State Management:** React Context + hooks (or Zustand)
421
+
422
+ **External APIs (Optional):**
423
+
424
+ - football-data.org for real-world data integration
425
+ - OpenFootball datasets for initial data bootstrapping
426
+
427
+ See the [Architecture Decision Record](docs/ADR-001-technology-stack.md) for detailed rationale.
428
+
429
+ See the [Game Design Document](docs/GDD.md) for:
430
+
431
+ - Complete feature specification
432
+ - Data models and ER diagrams
433
+ - UI wireframes and layout designs
434
+ - Performance requirements
435
+ - Development phases and milestones
436
+
437
+ ## 🚀 Getting Started
438
+
439
+ ### Prerequisites
440
+
441
+ - Node.js 18+ (LTS) and npm
442
+ - Git
443
+
444
+ ### Development Setup
445
+
446
+ ```bash
447
+ # Clone repository
448
+ git clone <repository-url>
449
+ cd football-management-sim
450
+
451
+ # Install dependencies
452
+ npm install
453
+
454
+ # Build the project
455
+ npm run build
456
+
457
+ # Run the application
458
+ npm start
459
+
460
+ # Run tests
461
+ npm test
462
+
463
+ # Lint code
464
+ npm run lint
465
+
466
+ # Format code
467
+ npm run format
468
+ ```
469
+
470
+ ### Project Structure
471
+
472
+ ```
473
+ .
474
+ ├── docs/
475
+ │ └── GDD.md # Complete Game Design Document
476
+ ├── src/ # Source code (to be created)
477
+ ├── tests/ # Test suite (to be created)
478
+ ├── data/ # Sample data, JSON fixtures
479
+ ├── TASKS.md # Development task tracking
480
+ ├── README.md # This file
481
+ └── .github/workflows/ # CI/CD pipelines
482
+ ```
483
+
484
+ ## 🧪 Testing
485
+
486
+ **Current Coverage: >89%** (Statements, Lines, Functions)
487
+
488
+ The project includes a comprehensive automated test suite:
489
+
490
+ - **Unit Tests:** Jest tests covering data models, utilities, match engine, tactics, transfer system, competitions, and UI components
491
+ - **Integration Tests:** End-to-end workflows for match simulation, transfer market, squad registration, and calendar
492
+ - **E2E Tests:** Cypress tests verifying complete user journeys through the React UI
493
+ - **Performance Benchmarks:** Automated benchmarks for match simulation speed and memory usage
494
+
495
+ ### Running Tests
496
+
497
+ ```bash
498
+ # Run all tests (unit + integration)
499
+ npm test
500
+
501
+ # Run with coverage report
502
+ npm run test:coverage
503
+
504
+ # Run E2E tests (Cypress)
505
+ npm run test:e2e
506
+
507
+ # Run all test suites
508
+ npm run test:all
509
+
510
+ # Run performance benchmarks
511
+ npm run benchmark
512
+ ```
513
+
514
+ ### Continuous Integration
515
+
516
+ Tests run automatically on every push and pull request via GitHub Actions:
517
+
518
+ - Node.js 18.x and 20.x matrix
519
+ - Linting, build, and test steps
520
+ - Coverage thresholds enforced
521
+ - Performance benchmarks included
522
+
523
+ See `.github/workflows/test.yml` for CI configuration.
524
+
525
+ ## 🏆 Beta Testing
526
+
527
+ The game has undergone structured beta testing with 5-10 external testers:
528
+
529
+ - **Performance:** Excellent - 27.2ms per match minute (target: <100ms)
530
+ - **Stability:** Zero crashes or data corruption
531
+ - **Memory:** 142MB steady-state (target: <500MB)
532
+ - **Feedback:** Positive overall, with minor bug fixes identified
533
+
534
+ See [Beta Test Report](BETA_TEST_REPORT.md) for full results.
535
+
536
+ ### Bug Reporting
537
+
538
+ Found an issue? Please use the [bug report template](.github/ISSUE_TEMPLATE/bug_report.md) to file a GitHub issue with detailed reproduction steps.
539
+
540
+ ### Performance Targets
541
+
542
+ - Match simulation: ≤100ms per match minute (achieved: 27.2ms)
543
+ - Memory usage: <500MB steady state (achieved: 142MB)
544
+ - Test coverage: ≥80% (achieved: 89.59%)
545
+ - UI frame rate: 60 FPS (achieved)
546
+
547
+ ## 🤝 Contributing
548
+
549
+ This is a solo development project for now. Contributions not accepted unless specified.
550
+
551
+ ## Saving and Loading
552
+
553
+ - Access **Save Game** from the HUD's quick actions or **Load Game** from the main menu.
554
+ - The game supports 10 manual save slots.
555
+ - Auto-save is enabled by default every 5 minutes (adjustable in Settings).
556
+
557
+ ## 📄 License
558
+
559
+ [To be determined]
560
+
561
+ ## 📚 Documentation
562
+
563
+ - [Game Design Document](docs/GDD.md) - Comprehensive design specification
564
+ - [Entity-Relationship Diagram](docs/ER-diagram.md) - Database schema visualization
565
+ - [TASKS.md](TASKS.md) - Development roadmap
566
+ - Inline code documentation (docstrings)
567
+
568
+ ## 🔮 Roadmap
569
+
570
+ 1. **Planning & Setup** (Tasks 1.1-1.4) → Complete GDD, choose tech stack, create data models
571
+ 2. **Core Game Systems** (Tasks 2.1-2.4) → Match engine, tactics, transfers, competitions
572
+ 3. **UI/UX & Polish** (Tasks 3.1-3.4) → Complete interface, match day UI, menus, assets
573
+ 4. **Testing & Deployment** (Tasks 4.1-4.4) → Automated tests, beta testing, performance tuning, release builds
574
+
575
+ ---
576
+
577
+ **Status:** Early development phase - No playable build available yet.
docs/readmes/hipstercheck-README.md ADDED
@@ -0,0 +1,531 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # hipstercheck
2
+ **Repository:** Julien-ser/hipstercheck
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-17T00:14:46Z
8
+
9
+ ---
10
+
11
+ # hipstercheck 🚀
12
+
13
+ **AI-Powered Code Review for Indie Developers**
14
+
15
+ > MVP READY FOR LAUNCH - Try the live demo at [https://hipstercheck.vercel.app](https://hipstercheck.vercel.app)
16
+
17
+ hipstercheck is an AI-powered code review tool that scans GitHub repositories for bugs, optimization suggestions, and best practices. Built for solo coders and small teams working with Python, ROS2, or ML frameworks, it delivers comprehensive reviews in under 60 seconds through a simple web interface.
18
+
19
+ ## Key Features
20
+
21
+ - 🔍 **Smart Code Analysis**: Detect bugs, performance issues, and style violations
22
+ - 🤖 **AI-Powered Reviews**: Fine-tuned open-source LLM for accurate suggestions
23
+ - 🚀 **GitHub Integration**: Seamless OAuth and repository scanning
24
+ - 📝 **Manual Upload & Paste**: Try the Quick Demo on the landing page - upload files or paste code snippets directly without GitHub login
25
+ - ⚡ **Fast Results**: Analysis completed in under 60 seconds
26
+ - 💰 **Affordable**: $10/month per user with free tier (1 repo scan/week)
27
+ - 🎯 **Specialized Support**: Python (PEP8), ROS2 best practices, ML frameworks (PyTorch, TensorFlow, scikit-learn)
28
+
29
+ ## 🎉 MVP Launch (March 2025)
30
+
31
+ hipstercheck is now live and ready for users!
32
+
33
+ **Live Demo**: [https://hipstercheck.vercel.app](https://hipstercheck.vercel.app)
34
+
35
+ ### What's Ready
36
+ - ✅ Full GitHub OAuth integration
37
+ - ✅ Repository scanning with rate limit handling
38
+ - ✅ AI-powered code review using fine-tuned Phi-2 model
39
+ - ✅ Support for Python, ROS2, ML frameworks (PyTorch, TensorFlow, scikit-learn)
40
+ - ✅ Manual file upload and code paste
41
+ - ✅ Dual-level caching (Redis/in-memory)
42
+ - ✅ Stripe payments (Pro tier $10/month)
43
+ - ✅ Free tier: 1 repo scan per week
44
+ - ✅ Deployed on Vercel (backend) + Streamlit Cloud (frontend)
45
+ - ✅ Comprehensive test suite passing
46
+ - ✅ Ready for community feedback
47
+
48
+ ### Join the Beta
49
+ We're looking for early adopters to test the MVP and provide feedback. Sign up for the free tier and help us improve!
50
+
51
+ 📢 **We're launching on Reddit and Indie Hackers!**
52
+ Follow our launch progress and share feedback:
53
+ - r/Startup_Ideas
54
+ - r/Python
55
+ - r/MLQuestions
56
+ - r/ROS2
57
+ - Indie Hackers
58
+
59
+ For launch materials and strategy, see [`launch/LAUNCH.md`](launch/LAUNCH.md).
60
+
61
+ ### Launch Metrics Goal
62
+ - 🎯 50 free tier signups in first 30 days
63
+ - 🎯 10 paid conversions
64
+ - 🎯 500+ total visitors
65
+ - 🎯 20+ meaningful feedback comments
66
+
67
+ ---
68
+
69
+ ## Quick Start
70
+
71
+ ### Prerequisites
72
+
73
+ - Python 3.9+
74
+ - Git
75
+ - GitHub account (for OAuth)
76
+ - Hugging Face account (for AI model, Phase 2)
77
+
78
+ ### 1. Configure GitHub OAuth App
79
+
80
+ Before running the app, you need to create a GitHub OAuth App:
81
+
82
+ 1. Go to [GitHub Developer Settings → OAuth Apps → New OAuth App](https://github.com/settings/developers)
83
+ 2. Fill in:
84
+ - **Application name**: hipstercheck (or your preferred name)
85
+ - **Homepage URL**: `http://localhost:8501` (or your deployed URL)
86
+ - **Authorization callback URL**: `http://localhost:8501` (or your deployed URL)
87
+ 3. Click "Register application"
88
+ 4. Copy the **Client ID** (generate a new client secret)
89
+ 5. Create a `.env` file from `.env.example` and add your credentials:
90
+ ```bash
91
+ cp .env.example .env
92
+ ```
93
+ 6. Edit `.env`:
94
+ ```bash
95
+ GITHUB_CLIENT_ID=your_client_id
96
+ GITHUB_CLIENT_SECRET=your_client_secret
97
+ APP_URL=http://localhost:8501
98
+ API_URL=http://localhost:8000
99
+ ```
100
+ *(Note: `APP_URL` must match the callback URL from step 2)*
101
+ *(Note: `API_URL` points to the FastAPI backend - start it with `python api.py`)*
102
+
103
+ ### 2. (Optional) Configure Stripe for Payments
104
+
105
+ To enable paid subscriptions ($10/month Pro plan):
106
+
107
+ 1. Create a [Stripe account](https://stripe.com) (if you don't have one)
108
+ 2. Get your API keys from Stripe Dashboard → Developers → API keys:
109
+ - **Secret key** (starts with `sk_test_` or `sk_live_`)
110
+ - **Publishable key** (starts with `pk_test_` or `pk_live_`)
111
+ 3. Create a product and price in Stripe Dashboard → Products:
112
+ - Create product: "hipstercheck Pro"
113
+ - Add price: $10/month (recurring)
114
+ - Copy the **Price ID** (starts with `price_`)
115
+ 4. Set up a webhook endpoint in Stripe Dashboard → Developers → Webhooks:
116
+ - Endpoint URL: `https://your-app.com/api/stripe/webhook` (or `http://localhost:8000/stripe/webhook` for local)
117
+ - Select events: `checkout.session.completed`, `customer.subscription.updated`, `customer.subscription.deleted`
118
+ - Copy the **Webhook secret** (starts with `whsec_`)
119
+ 5. Add the following to your `.env`:
120
+
121
+ ```bash
122
+ # Stripe API keys
123
+ STRIPE_SECRET_KEY=sk_test_your_secret_key
124
+ STRIPE_PUBLIC_KEY=pk_test_your_publishable_key
125
+ STRIPE_WEBHOOK_SECRET=whsec_your_webhook_secret
126
+ STRIPE_PRICE_ID=price_your_price_id
127
+ ```
128
+
129
+ **Note**: For local testing, you can use Stripe test mode and test cards (e.g., `4242 4242 4242 4242`). The fallback price ID in the code is for demo purposes; you should create your own.
130
+
131
+ ### 3. Install Dependencies
132
+
133
+ ```bash
134
+ pip install -r requirements.txt
135
+ ```
136
+
137
+ ### 5. (Optional) Configure Redis for Caching
138
+
139
+ To enable persistent caching across app restarts and multiple users:
140
+
141
+ 1. Install and run Redis:
142
+ ```bash
143
+ # Ubuntu/Debian
144
+ sudo apt-get install redis-server
145
+ sudo systemctl start redis
146
+
147
+ # macOS
148
+ brew install redis
149
+ brew services start redis
150
+
151
+ # Docker
152
+ docker run -p 6379:6379 redis:alpine
153
+ ```
154
+
155
+ 2. Add to your `.env`:
156
+ ```bash
157
+ REDIS_URL=redis://localhost:6379/0
158
+ ```
159
+
160
+ If Redis is not available, the app will fall back to in-memory caching.
161
+
162
+ ### 6. (Optional) Configure Rate Limit Buffer
163
+
164
+ The app automatically monitors GitHub API rate limits and will wait when approaching the limit. You can adjust the buffer:
165
+
166
+ ```bash
167
+ RATE_LIMIT_BUFFER=10
168
+ ```
169
+
170
+ This sets how many remaining requests trigger a wait before the reset time.
171
+
172
+ ```bash
173
+ streamlit run streamlit_app.py
174
+ ```
175
+
176
+ ### 3. (Optional) Run FastAPI Backend
177
+
178
+ The FastAPI microservice provides the code analysis endpoint. To run it separately:
179
+
180
+ ```bash
181
+ # Start the API server
182
+ python api.py
183
+
184
+ # Or with uvicorn directly
185
+ uvicorn api:app --host 0.0.0.0 --port 8000 --reload
186
+ ```
187
+
188
+ The API will be available at `http://localhost:8000` with:
189
+ - `GET /health` - Health check endpoint
190
+ - `POST /analyze` - Analyze single code snippet (5s timeout)
191
+ - `POST /analyze/batch` - Analyze multiple snippets (max 50)
192
+
193
+ ### 4. Authenticate & Use
194
+
195
+ 1. Click **"Login with GitHub"** in the app
196
+ 2. Authorize the OAuth app on GitHub
197
+ 3. View your repositories in the dashboard
198
+ 4. (Future) Select a repo to analyze code
199
+
200
+ ## Backend Deployment
201
+
202
+ ### Backend Deployment
203
+
204
+ The FastAPI backend is configured for serverless deployment on Vercel using Mangum.
205
+
206
+ 1. **Install Vercel CLI** (optional but recommended):
207
+ ```bash
208
+ npm i -g vercel
209
+ ```
210
+
211
+ 2. **Login to Vercel**:
212
+ ```bash
213
+ vercel login
214
+ ```
215
+
216
+ 3. **Link your project**:
217
+ ```bash
218
+ vercel link
219
+ ```
220
+ Follow the prompts to link your project. This uses the existing `vercel.json` configuration.
221
+
222
+ 4. **Set environment variables in Vercel dashboard**:
223
+
224
+ After linking, go to your project settings in the Vercel dashboard (https://vercel.com/[your-org]/hipstercheck/settings/environment-variables) and add all variables from `.env.example`:
225
+
226
+ - `GITHUB_CLIENT_ID` - From your GitHub OAuth app
227
+ - `GITHUB_CLIENT_SECRET` - From your GitHub OAuth app
228
+ - `APP_URL` - Your Vercel frontend URL (e.g., `https://hipstercheck.vercel.app`)
229
+ - `HF_TOKEN` - Your Hugging Face token (for model inference)
230
+ - `STRIPE_SECRET_KEY` - Your Stripe secret key
231
+ - `STRIPE_PUBLIC_KEY` - Your Stripe publishable key
232
+ - `STRIPE_WEBHOOK_SECRET` - Your Stripe webhook signing secret
233
+ - `STRIPE_PRICE_ID` - Your Stripe price ID for Pro subscription
234
+ - `REDIS_URL` (optional) - Redis connection URL for persistent caching
235
+ - `CACHE_TTL_HOURS` (optional) - Cache TTL in hours, default 24
236
+ - `RATE_LIMIT_BUFFER` (optional) - GitHub API rate limit buffer, default 10
237
+
238
+ Alternatively, set them via CLI:
239
+ ```bash
240
+ vercel env add GITHUB_CLIENT_ID production
241
+ vercel env add HF_TOKEN production
242
+ # ... repeat for each variable
243
+ ```
244
+
245
+ 5. **Deploy the backend**:
246
+ ```bash
247
+ vercel --prod
248
+ ```
249
+
250
+ The API will be available at:
251
+ - `https://your-project.vercel.app/api/health` (health check)
252
+ - `https://your-project.vercel.app/api/analyze` (analysis endpoint)
253
+ - `https://your-project.vercel.app/api/stripe/*` (Stripe endpoints)
254
+
255
+ 6. **Enable Vercel Analytics**:
256
+
257
+ In the Vercel dashboard, enable **Analytics** for your project to monitor:
258
+ - Request count and duration
259
+ - Error rates
260
+ - Function invocations
261
+
262
+ Vercel automatically collects backend metrics. View them under the "Analytics" tab.
263
+
264
+ ### Updating Frontend Configuration
265
+
266
+ After deploying the backend, update your `.env` file for the Streamlit frontend:
267
+
268
+ ```bash
269
+ API_URL=https://your-project.vercel.app
270
+ ```
271
+
272
+ ### Monitoring & Logs
273
+
274
+ - **Real-time logs**: `vercel logs your-project.vercel.app --follow`
275
+ - **Performance metrics**: Vercel dashboard → Analytics
276
+ - **Alerts**: Set up notifications in Vercel dashboard
277
+
278
+ ### Notes
279
+
280
+ - Backend uses `mangum` for ASGI-to-serverless adapter
281
+ - Cold start: ~1-2 seconds for first request after inactivity
282
+ - Redis recommended for production caching; falls back to in-memory if unavailable
283
+ - Ensure Hugging Face token has `read` permissions for model access
284
+
285
+ ---
286
+
287
+ ## Frontend Deployment
288
+
289
+ The Streamlit frontend can be deployed either on Streamlit Community Cloud (recommended) or on Vercel using Docker.
290
+
291
+ ### Prerequisites
292
+
293
+ - GitHub repository with the hipstercheck code
294
+ - Backend already deployed on Vercel (or another accessible URL)
295
+ - Environment variables configured (see `.env.example`)
296
+
297
+ ### Option 1: Streamlit Community Cloud (Recommended)
298
+
299
+ Streamlit Community Cloud provides free hosting for Streamlit apps with minimal configuration.
300
+
301
+ 1. **Push your code to GitHub** (if not already):
302
+ ```bash
303
+ git add .
304
+ git commit -m "Prepare for deployment"
305
+ git push origin main
306
+ ```
307
+
308
+ 2. **Go to [Streamlit Cloud](https://streamlit.io/cloud) and sign up** (you can use your GitHub account).
309
+
310
+ 3. **Create a new app**:
311
+ - Click **"New app"**.
312
+ - Choose your GitHub repository and branch (e.g., `main`).
313
+ - Set **Main file path** to `streamlit_app.py`.
314
+ - Leave the other settings as default.
315
+
316
+ 4. **Configure environment variables**:
317
+ In the app's settings (gear icon), go to **Secrets** and add each variable from your `.env.example`:
318
+ - `GITHUB_CLIENT_ID`
319
+ - `GITHUB_CLIENT_SECRET`
320
+ - `APP_URL` (set to your Streamlit Cloud URL, e.g., `https://your-app.streamlit.app`)
321
+ - `API_URL` (set to your deployed backend URL, e.g., `https://hipstercheck.vercel.app`)
322
+ - `HF_TOKEN` (if using local model inference)
323
+ - `STRIPE_SECRET_KEY`, `STRIPE_PUBLIC_KEY`, `STRIPE_WEBHOOK_SECRET`, `STRIPE_PRICE_ID` (if enabling payments)
324
+ - Optional: `REDIS_URL`, `CACHE_TTL_HOURS`, `RATE_LIMIT_BUFFER`
325
+
326
+ Alternatively, you can set these in the **Settings → Secrets** section.
327
+
328
+ 5. **Deploy**:
329
+ Click **"Deploy"**. Streamlit Cloud will build and deploy your app automatically. The deployment usually takes a few minutes.
330
+
331
+ 6. **Verify**:
332
+ Once deployed, open your app URL and test the functionality. Ensure the OAuth callback works and that code analysis calls reach the backend.
333
+
334
+ ### Option 2: Vercel with Docker (Advanced)
335
+
336
+ If you prefer to host the frontend on Vercel, you can containerize the Streamlit app using Docker.
337
+
338
+ 1. **Create a `Dockerfile`** in the project root (if not already present):
339
+ ```dockerfile
340
+ FROM python:3.11-slim
341
+
342
+ WORKDIR /app
343
+
344
+ RUN apt-get update && apt-get install -y --no-install-recommends \
345
+ build-essential \
346
+ && rm -rf /var/lib/apt/lists/*
347
+
348
+ COPY requirements.txt .
349
+ RUN pip install --no-cache-dir -r requirements.txt
350
+
351
+ COPY . .
352
+
353
+ ENV PYTHONUNBUFFERED=1
354
+ EXPOSE 8501
355
+
356
+ CMD ["sh", "-c", "uvicorn api:app --host 0.0.0.0 --port 8000 & streamlit run streamlit_app.py --server.port $PORT --server.address 0.0.0.0"]
357
+ ```
358
+
359
+ This Dockerfile runs both the FastAPI backend (on port 8000 internally) and the Streamlit frontend (on the Vercel-provided `$PORT`). The frontend will communicate with the backend via `http://localhost:8000`.
360
+
361
+ 2. **Update `vercel.json`** to use Docker:
362
+ ```json
363
+ {
364
+ "version": 2,
365
+ "builds": [
366
+ {
367
+ "src": "Dockerfile",
368
+ "use": "@vercel/docker"
369
+ }
370
+ ],
371
+ "routes": [
372
+ {
373
+ "src": "/(.*)",
374
+ "dest": "/"
375
+ }
376
+ ]
377
+ }
378
+ ```
379
+
380
+ 3. **Set environment variables in Vercel**:
381
+ When you create a new Vercel project (or update the existing one), add all environment variables from `.env.example` in the Vercel dashboard → Project Settings → Environment Variables. Ensure `APP_URL` matches your Vercel frontend URL and `API_URL` is set to `http://localhost:8000` (since the backend is in the same container).
382
+
383
+ 4. **Deploy**:
384
+ ```bash
385
+ vercel --prod
386
+ ```
387
+
388
+ 5. **Notes**:
389
+ - The combined container starts both the backend and frontend. FastAPI runs in the background and is only accessible internally; all external traffic goes to Streamlit.
390
+ - If you already have a separate backend deployment, you can modify `API_URL` to point to that external URL instead.
391
+ - Vercel's Docker builds may take longer than serverless builds.
392
+
393
+ ### Adding a Custom Domain
394
+
395
+ Both Streamlit Cloud and Vercel support custom domains:
396
+
397
+ - **Streamlit Cloud**: In your app settings, go to **Settings → Custom domain**, add your domain, and follow the DNS verification instructions.
398
+ - **Vercel**: In the Vercel dashboard, go to your project's **Domains** settings and add your domain. Vercel automatically provisions an SSL certificate.
399
+
400
+ After adding a custom domain, update the `APP_URL` environment variable to match the new domain.
401
+
402
+ ---
403
+
404
+ **Important**: The frontend requires the backend to be running and accessible. Ensure your backend deployment is healthy before connecting the frontend.
405
+
406
+ ## Project Structure
407
+
408
+ ```
409
+ hipstercheck/
410
+ ├── api.py # FastAPI microservice (Phase 3)
411
+ ├── streamlit_app.py # Main Streamlit frontend
412
+ ├── requirements.txt # Python dependencies
413
+ ├── .env # Environment variables (create from .env.example)
414
+ ├── dataset/ # Code review dataset collection
415
+ │ ├── data_collector.py # Dataset collection script
416
+ │ ├── code_reviews.jsonl # Combined dataset (generated)
417
+ │ ├── split_train.jsonl # Training split
418
+ │ ├── split_val.jsonl # Validation split
419
+ │ ├── split_test.jsonl # Test split
420
+ │ └── README.md # Dataset documentation
421
+ ├── models/ # Fine-tuned AI model
422
+ ├── prompts/ # Prompt templates for different languages
423
+ ├── tests/ # Unit and integration tests
424
+ ├── README.md # This file
425
+ └── TASKS.md # Development task tracking
426
+ ```
427
+
428
+ ## 📋 Development Status
429
+
430
+ **All Phases Complete** ✅ MVP Ready for Launch
431
+
432
+ ### Phase 1: Foundation & GitHub Integration ✅ Completed
433
+ - [x] Initialize Streamlit project structure
434
+ - [x] Implement GitHub OAuth authentication
435
+ - [x] Build repo scanning engine
436
+ - [x] Configure GitHub API rate limiting with caching
437
+
438
+ ### Phase 2: Model Training & Code Analysis ✅ Completed
439
+ - [x] Collect code review dataset (15K+ examples)
440
+ - [x] Select base LLM (Microsoft Phi-2, 2.7B parameters)
441
+ - [x] Fine-tune model with LoRA on code review generation
442
+ - [x] Create prompt engineering templates for Python, ROS2, ML frameworks
443
+
444
+ ### Phase 3: App Integration & Deployment Prep ✅ Completed
445
+ - [x] Wrap model in FastAPI microservice
446
+ - [x] Integrate model calls into Streamlit with color-coded UI
447
+ - [x] Implement dual-level caching (repo scan + review results)
448
+ - [x] Set up Stripe subscription ($10/month Pro tier)
449
+
450
+ ### Phase 4: Testing, Deployment & Validation ✅ Completed
451
+ - [x] Test with personal ROS2/ML projects
452
+ - [x] Deploy backend on Vercel (serverless FastAPI)
453
+ - [x] Deploy Streamlit frontend on Streamlit Community Cloud
454
+ - [x] Validate via Reddit/Indie Hackers (launch materials prepared)
455
+
456
+ **Next**: Community feedback collection and iterative improvements based on user input. See [`launch/ROADMAP.md`](launch/ROADMAP.md) for post-launch plan.
457
+
458
+ ## Caching
459
+
460
+ hipstercheck implements **two-level caching** to optimize performance and reduce costs:
461
+
462
+ ### 1. Repository Scan Cache
463
+ Scanning a repository clones it and extracts the file tree, which can be slow and rate-limited by GitHub. Results are cached for 24 hours with Redis (preferred) or in-memory fallback.
464
+
465
+ **Cache Key**: Repository name + branch (hashed)
466
+
467
+ **Cache Invalidation**:
468
+ - Automatic: Entries expire after 24 hours
469
+ - Manual: Use "Re-Scan" button for fresh scan
470
+ - Global: "Clear All Cache" button clears everything
471
+
472
+ **Configuration**:
473
+ ```bash
474
+ # Add to .env (optional)
475
+ REDIS_URL=redis://localhost:6379/0
476
+ CACHE_TTL_HOURS=24
477
+ ```
478
+
479
+ If Redis is unavailable, the app falls back to in-memory caching (lost on restart).
480
+
481
+ ### 2. Code Review Cache
482
+ Individual code file analyses are cached to avoid unnecessary model inference. This significantly reduces compute costs and improves response time.
483
+
484
+ **Cache Key**: SHA256 hash of (code content + language)
485
+
486
+ **Statistics**: Real-time hit/miss rates displayed in the UI. Typical hit rates >60% for repeated analyses.
487
+
488
+ **Configuration**:
489
+ ```bash
490
+ CACHE_TTL_HOURS=24 # Review cache TTL (default: 24 hours)
491
+ ```
492
+
493
+ ### Cache Performance Tips
494
+ - ✅ **High cache hit rates** when scanning the same repo multiple times
495
+ - ✅ **Fast re-scans** - bypass GitHub clone, instant results from cache
496
+ - ✅ **Reduced model calls** - saves GPU/compute costs
497
+ - ✅ **Offline capability** - memory cache works without Redis
498
+
499
+ ### Monitoring
500
+ The dashboard displays:
501
+ - **Scan Cache**: Backend type (Redis/Memory), total cached repos
502
+ - **Review Cache**: Hit rate, total requests, backend
503
+ - **Expandable details** in the "📊 Detailed Cache Statistics" section
504
+
505
+ Use "Clear All Cache" to reset statistics and force fresh analyses.
506
+
507
+ ## Architecture
508
+
509
+ ### Tech Stack
510
+ - **Frontend**: Streamlit (responsive UI, rapid prototyping)
511
+ - **Backend**: FastAPI (async API, high performance)
512
+ - **AI Model**: Phi-2 (2.7B parameters) fine-tuned with LoRA for code review
513
+ - **Training**: PyTorch + Hugging Face Transformers + PEFT (LoRA)
514
+ - **Database**: SQLite/Redis for caching
515
+ - **Deployment**: Vercel (serverless)
516
+ - **Payment**: Stripe Checkout
517
+
518
+ ### Data Flow
519
+ 1. User authenticates via GitHub OAuth
520
+ 2. Repository files are scanned and parsed
521
+ 3. AI model analyzes code for issues
522
+ 4. Results displayed with severity levels and suggestions
523
+ 5. Cached for future reference
524
+
525
+ ## License
526
+
527
+ MIT License - see LICENSE file for details
528
+
529
+ ## Contributing
530
+
531
+ This is a solo project for now. For feedback or collaboration, reach out via GitHub Issues.
docs/readmes/internlearningnetwork-README.md ADDED
@@ -0,0 +1,327 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # internlearningnetwork
2
+ **Repository:** Julien-ser/internlearningnetwork
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** HTML
7
+ **Last Updated:** 2026-03-16T22:17:58Z
8
+
9
+ ---
10
+
11
+ # InternLearningNetwork
12
+
13
+ **Mission:** Allows interns all over the place to share anything they found/learned in a new blog-like system, with a gamified way of levelling up with new skills and points and also points for sharing something that gives other users skills.
14
+
15
+ ## Tech Stack
16
+
17
+ - **Frontend**: React + Vite, TypeScript
18
+ - **Backend**: Node.js + Express, TypeScript
19
+ - **Database**: PostgreSQL
20
+ - **ORM**: Prisma
21
+ - **Authentication**: JWT + bcrypt
22
+ - **Architecture**: Monorepo structure
23
+
24
+ ## Project Structure
25
+
26
+ ```
27
+ internlearningnetwork/
28
+ ├── client/ # React frontend application
29
+ │ ├── src/
30
+ │ │ ├── components/ # Reusable UI components
31
+ │ │ ├── pages/ # Page components
32
+ │ │ ├── services/ # API client
33
+ │ │ └── types/ # TypeScript definitions
34
+ │ └── package.json
35
+ ├── server/ # Express backend API
36
+ │ ├── src/
37
+ │ │ ├── routes/ # API route handlers
38
+ │ │ ├── middleware/ # Auth, validation, etc.
39
+ │ │ ├── services/ # Business logic
40
+ │ │ ├── prisma/ # Database schema & migrations
41
+ │ │ └── utils/ # Helper functions
42
+ │ └── package.json
43
+ ├── shared/ # Shared TypeScript types & utilities
44
+ │ └── src/
45
+ │ ├── types/ # Common type definitions
46
+ │ └── validations/ # Shared validation schemas
47
+ ├── .github/workflows/
48
+ │ └── test.yml # CI/CD pipeline (to be created)
49
+ ├── TASKS.md # Development task list
50
+ └── README.md # This file
51
+ ```
52
+
53
+ ## Architecture Diagram
54
+
55
+ ```
56
+ ┌─────────────────────────────────────────────────────────────┐
57
+ │ Frontend (React) │
58
+ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────┐ │
59
+ │ │ PostFeed │ │CreatePost │ │ UserProfile │ │
60
+ │ └─────────────┘ └─────────────┘ └─────────────────────┘ │
61
+ │ │ │ │ │
62
+ │ └────────────────┼───────────────────┘ │
63
+ │ │ │
64
+ │ ┌───────▼────────┐ │
65
+ │ │ API Client │ (axios/fetch) │
66
+ │ └────────┬───────┘ │
67
+ └───────────────────────────┼─────────────────────────────────┘
68
+ │ HTTPS/REST API
69
+ │ JWT Auth
70
+ ┌───────────────────────────▼─────────────────────────────────┐
71
+ │ Backend (Express) │
72
+ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────┐ │
73
+ │ │ Auth API │ │ Posts API │ │ Skills & Points │ │
74
+ │ └─────────────┘ └─────────────┘ └─────────────────────┘ │
75
+ │ │ │ │ │
76
+ │ └────────────────┼───────────────────┘ │
77
+ │ │ │
78
+ │ ┌───────▼────────┐ │
79
+ │ │ Middleware │ (auth, validation) │
80
+ │ └────────┬───────┘ │
81
+ │ │ │
82
+ │ ┌───────▼────────┐ │
83
+ │ │ Services │ (business logic) │
84
+ │ └────────┬───────┘ │
85
+ └─────────���─────────────────┼─────────────────────────────────┘
86
+ │ Prisma ORM
87
+ ┌───────────────────────────▼─────────────────────────────────┐
88
+ │ PostgreSQL Database │
89
+ │ ┌────────┐ ┌───────┐ ┌──────┐ ┌─────────┐ ┌──────────┐ │
90
+ │ │ users │ │ posts │ │skills││user_skills│points_log│ │
91
+ │ └────────┘ └───────┘ └──────┘ └─────────┘ └──────────┘ │
92
+ │ ┌────────┐ ┌──────────────────────────────────────────┐ │
93
+ │ │ levels │ │ levels (calculated/derived) │ │
94
+ │ └────────┘ └──────────────────────────────────────────┘ │
95
+ └─────────────────────────────────────────────────────────────┘
96
+ ```
97
+
98
+ ## Database Schema
99
+
100
+ ### Core Tables
101
+ - **users**: id, email, password_hash, username, created_at, updated_at
102
+ - **posts**: id, title, content, author_id (FK), created_at, updated_at
103
+ - **skills**: id, name, description, created_at
104
+ - **post_skills**: post_id (FK), skill_id (FK) - junction table for many-to-many
105
+ - **user_skills**: user_id (FK), skill_id (FK), claimed_at, source_post_id (FK) - tracks skills users have claimed
106
+ - **points_log**: id, user_id (FK), points, action_type, reference_id (post/skill), created_at
107
+ - **levels**: id, level_number, min_points, max_points (or use calculated levels)
108
+
109
+ ### Relationships
110
+ - User ↔ Post: One-to-many (one author, many posts)
111
+ - Post ↔ Skill: Many-to-many via post_skills
112
+ - User ↔ Skill: Many-to-many via user_skills (skills they've claimed)
113
+ - User ↔ Points: One-to-many via points_log (point history)
114
+ - User ↔ Level: Calculated based on total points
115
+
116
+ ## Current Progress
117
+
118
+ **Phase 1: Planning & Setup**
119
+ - ✅ Technical stack defined (Node.js + Express, React, PostgreSQL, Prisma, JWT)
120
+ - ✅ Architecture diagram created
121
+ - ✅ Monorepo structure initialized with package.json files
122
+ - ✅ Root workspace configuration created
123
+ - ✅ Prisma schema designed
124
+ - ✅ `.env.example` created and dependencies installed
125
+
126
+ **Phase 2: Core Backend & Authentication**
127
+ - ✅ Implemented user registration/login endpoints with JWT token generation using bcrypt
128
+ - `POST /api/auth/register` - User registration with email, username, password
129
+ - `POST /api/auth/login` - User login, returns JWT token
130
+ - `GET /api/auth/me` - Get authenticated user's profile (requires Bearer token)
131
+ - ✅ Created CRUD API for blog posts with validation middleware
132
+ - Posts include skill_tags array for associating skills
133
+ - ✅ Built skill management system with full CRUD endpoints (authenticated)
134
+ - `GET /api/skills` - List all skills
135
+ - `GET /api/skills/:id` - Get skill details
136
+ - `POST /api/skills` - Create new skill (requires authentication)
137
+ - `PUT /api/skills/:id` - Update skill (requires authentication)
138
+ - `DELETE /api/skills/:id` - Delete skill (requires authentication)
139
+ - Skills can be associated with posts via skill_tags field
140
+ - ✅ Implemented post approval system that assigns skills to authors
141
+ - `PUT /api/posts/:id/approve` - Approve post and automatically assign associated skills to the author (requires authentication)
142
+
143
+ ## Getting Started
144
+
145
+ ### Prerequisites
146
+ - Node.js 18+ installed
147
+ - PostgreSQL database running
148
+ - Git
149
+
150
+ ### 1. Environment Configuration
151
+
152
+ Copy the example environment file and update with your values:
153
+
154
+ ```bash
155
+ cp .env.example .env
156
+ ```
157
+
158
+ Edit `.env` and configure:
159
+ - `DATABASE_URL`: Your PostgreSQL connection string
160
+ - `JWT_SECRET`: A secure random string for JWT signing
161
+
162
+ ### 2. Database Setup
163
+
164
+ Generate Prisma client and create/update database schema:
165
+
166
+ ```bash
167
+ npx prisma generate
168
+ npx prisma db push
169
+ ```
170
+
171
+ (Optional) Seed the database with initial data:
172
+
173
+ ```bash
174
+ npx prisma db seed
175
+ ```
176
+
177
+ ### 3. Running the Application
178
+
179
+ **Development mode** (runs both frontend and backend):
180
+
181
+ ```bash
182
+ npm run dev
183
+ ```
184
+
185
+ Or run them separately:
186
+
187
+ ```bash
188
+ # Backend only (server on port 3001)
189
+ npm run dev:server
190
+
191
+ # Frontend only (client on port 5173)
192
+ npm run dev:client
193
+ ```
194
+
195
+ **Build for production:**
196
+
197
+ ```bash
198
+ npm run build
199
+ ```
200
+
201
+ **Run linter:**
202
+
203
+ ```bash
204
+ npm run lint
205
+ ```
206
+
207
+ ### 4. Access the Application
208
+
209
+ - Frontend: http://localhost:5173
210
+ - Backend API: http://localhost:3001/api
211
+ - API Health check: http://localhost:3001/api/health (to be implemented)
212
+
213
+ ## Deployment
214
+
215
+ ### Prerequisites
216
+ - Accounts on Vercel/Netlify (for frontend) and Railway/Render (for backend)
217
+ - A PostgreSQL database (e.g., Supabase, Neon, AWS RDS, or any PostgreSQL provider)
218
+ - GitHub repository connected to your deployment platforms
219
+
220
+ ### Frontend Deployment (Vercel / Netlify)
221
+
222
+ 1. Import your repository as a new project in Vercel or Netlify.
223
+ 2. Set the **Root Directory** to `client` (since it's a monorepo).
224
+ 3. Configure build settings:
225
+ - **Build Command**: `npm run build` (or `npm run build --workspace=client`)
226
+ - **Output Directory**: `dist` (Vite's default)
227
+ 4. Add environment variables:
228
+ - `VITE_API_URL`: Your backend API URL (e.g., https://your-backend.onrender.com/api)
229
+ 5. Deploy. The platform will install dependencies, build, and serve the static files.
230
+
231
+ ### Backend Deployment (Railway / Render)
232
+
233
+ 1. Create a new service on Railway or Render and connect your repository.
234
+ 2. Set the **Root Directory** to `server`.
235
+ 3. Configure build and start commands:
236
+ - **Build Command**: `npm run build` (or `npm run build --workspace=server`)
237
+ - **Start Command**: `npm start` (or `npm start --workspace=server`)
238
+ 4. Add environment variables:
239
+ - `DATABASE_URL`: Your PostgreSQL connection string
240
+ - `JWT_SECRET`: A secure random string for JWT signing
241
+ - `NODE_ENV`: `production`
242
+ - (Optional) `PORT`: The port your server should listen on (default 3001)
243
+ 5. Add a **Post-Deploy** script (Railway) or **After Build** hook to run migrations and seed:
244
+ - `npx prisma migrate deploy`
245
+ - `npx prisma db seed` (optional, for demo data)
246
+ 6. Deploy. The platform will build the TypeScript code, apply database migrations, and start the server.
247
+
248
+ ### Database Setup
249
+ - Create a PostgreSQL database and obtain the connection URL.
250
+ - Ensure the database user has privileges to create tables and indexes.
251
+ - Update the `DATABASE_URL` in your production environment.
252
+
253
+ ### Notes
254
+ - The project uses Prisma ORM. Migrations are stored in `server/prisma/migrations`.
255
+ - The seed script (`server/prisma/seed.ts`) creates demo skills, levels, and an admin user (`admin@example.com` / `admin123`).
256
+ - For security, change the default admin password after first login and use a strong `JWT_SECRET`.
257
+
258
+ ## API Endpoints
259
+
260
+ ### Authentication ✅
261
+ - `POST /api/auth/register` - User registration
262
+ - Body: `{ email, username, password }`
263
+ - Returns: `{ message, user: { id, email, username, createdAt }, token }`
264
+ - `POST /api/auth/login` - User login
265
+ - Body: `{ email, password }`
266
+ - Returns: `{ message, user: { id, email, username, createdAt, totalPoints, level }, token }`
267
+ - `GET /api/auth/me` - Get current user profile **(requires Bearer token)**
268
+ - Headers: `Authorization: Bearer <jwt_token>`
269
+ - Returns: `{ user: { id, email, username, createdAt, totalPoints, level, userSkills } }`
270
+
271
+ ### Posts
272
+ - `GET /api/posts` - List all posts (with skill tags)
273
+ - `GET /api/posts/:id` - Get single post
274
+ - `POST /api/posts` - Create new post
275
+ - `PUT /api/posts/:id` - Update post
276
+ - `DELETE /api/posts/:id` - Delete post
277
+
278
+ ### Skills
279
+ - `GET /api/skills` - List all skills
280
+ - `GET /api/skills/:id` - Get skill details
281
+ - `POST /api/skills` - Create new skill (requires authentication)
282
+ - `PUT /api/skills/:id` - Update skill (requires authentication)
283
+ - `DELETE /api/skills/:id` - Delete skill (requires authentication)
284
+
285
+ ### Claims (Skill Claiming)
286
+ - `POST /api/claims/posts/:postId/skills/:skillId/claim` - Claim a skill from a post (requires authentication)
287
+ - Awards 5 points to the post author
288
+ - Adds the skill to the claimant's profile
289
+ - `GET /api/claims/user/skills` - Get authenticated user's claimed skills (requires authentication)
290
+
291
+ ### Leveling
292
+ - `GET /api/level` - Get user's current level information (requires `userId` query param)
293
+ - Query: `?userId=1`
294
+ - Returns: `{ userId, totalPoints, currentLevel, levelName, minPoints, maxPoints, pointsToNextLevel }`
295
+ - `POST /api/level/calculate` - Calculate what level a given point total would be
296
+ - Body: `{ points: 250 }`
297
+ - Returns: `{ points, level, levelName, minPoints, maxPoints, pointsToNextLevel }`
298
+ - `POST /api/level/update` - Update user's level based on current points (called after point allocations)
299
+ - Query: `?userId=1`
300
+ - Returns: `{ userId, totalPoints, currentLevel, levelName, message }`
301
+
302
+ **Leveling Algorithm:** Exponential thresholds with base 100 and multiplier 1.5
303
+ - Level 1: 0 - 99 points
304
+ - Level 2: 100 - 149 points
305
+ - Level 3: 150 - 224 points
306
+ - Level 4: 225 - 336 points
307
+ - Level 5: 337 - 505 points
308
+ - Level 6: 506 - 758 points
309
+ - Level 7: 759 - 1138 points
310
+ - Level 8: 1139 - 1707 points
311
+ - Level 9: 1708 - 2561 points
312
+ - Level 10: 2562+ points
313
+
314
+ ### Posts
315
+ - Skills can be attached to posts (like hashtags)
316
+ - Other users can "claim" skills from posts
317
+ - Claiming a skill adds it to the user's profile AND awards points to the post author
318
+ 4. **Gamification**:
319
+ - Points awarded: +10 for creating a post, +5 per skill claimed from your post
320
+ - Level progression based on total points
321
+ - Points log tracks all transactions
322
+ 5. **User Profiles**: Display level, points, earned skills with dates, recent activity
323
+ 6. **Leaderboard**: Competitive ranking based on points and level
324
+
325
+ ## License
326
+
327
+ TBD
docs/readmes/invoice-resolver-ai-README.md ADDED
@@ -0,0 +1,427 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # invoice-resolver-ai
2
+ **Repository:** Julien-ser/invoice-resolver-ai
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-21T03:29:21Z
8
+
9
+ ---
10
+
11
+ # Invoice Resolver AI
12
+
13
+ **An autonomous AI agent for resolving invoice disputes and recovering unpaid invoices.**
14
+
15
+ ---
16
+
17
+ ## Problem
18
+
19
+ Freelancers and small SaaS founders waste hours chasing unpaid invoices, drafting dispute letters, and navigating payment platforms. Payment delays and chargebacks are common, but resolution is manual and frustrating.
20
+
21
+ ## Solution
22
+
23
+ An intelligent AI agent that monitors Stripe/PayPal/Bank feeds, detects overdue or disputed invoices, and automatically:
24
+
25
+ - Sends polite follow-ups with payment links
26
+ - Drafts formal dispute letters (with evidence attachments) for chargebacks
27
+ - Escalates to small claims paperwork if needed (pre-filled forms)
28
+ - Learns which templates and timing work best (A/B testing)
29
+
30
+ ## Features
31
+
32
+ - **Multi-provider Integration**: Stripe, PayPal, and Plaid (bank feeds)
33
+ - **AI-Powered Dispute Drafting**: GPT-4/Claude integration for legal letters
34
+ - **Automated Email Campaigns**: Follow-up sequences (3, 7, 14 days)
35
+ - **PDF Generation**: Dispute letters and small claims forms
36
+ - **A/B Testing**: Optimize templates and timing
37
+ - **Analytics Dashboard**: Recovery rates, campaign performance
38
+ - **Admin Panel**: User management, system health monitoring
39
+ - **Freemium Model**: 5 free invoices/month, $19/mo unlimited
40
+
41
+ ## Tech Stack
42
+
43
+ - **Backend**: Python 3.11-3.13, FastAPI, PostgreSQL (Python 3.14 not yet supported due to pydantic-core compatibility)
44
+ - **Task Queue**: Celery + Redis
45
+ - **AI Integration**: OpenAI GPT-4 / Anthropic Claude
46
+ - **Payment APIs**: Stripe, PayPal REST SDK, Plaid
47
+ - **Email**: SendGrid / SMTP with Jinja2 templates
48
+ - **Dashboard**: Streamlit
49
+ - **PDF**: WeasyPrint or ReportLab
50
+ - **Deployment**: Docker, GitHub Actions CI/CD
51
+
52
+ ## Project Structure
53
+
54
+ ```
55
+ .
56
+ ├── README.md # Project documentation
57
+ ├── TASKS.md # Development task list (track progress)
58
+ ├── schema.sql # PostgreSQL database schema
59
+ ├── docs/
60
+ │ └── database-er-diagram.md # ER diagram and table docs
61
+ ├── .github/
62
+ │ └── workflows/
63
+ │ └── test.yml # CI pipeline (to be created)
64
+ ├── src/ # Source code (to be created)
65
+ │ ├── api/ # FastAPI endpoints
66
+ │ ├── core/ # Configuration, logging
67
+ │ ├── models/ # SQLAlchemy models
68
+ │ ├── integrations/ # Stripe/PayPal/Plaid clients
69
+ │ ├── email/ # Email automation
70
+ │ ├── ai/ # AI dispute drafter
71
+ │ ├── pdf/ # PDF generator
72
+ │ ├── dashboard/ # Streamlit app
73
+ │ ├── ab_testing/ # A/B testing framework
74
+ │ ├── admin/ # Admin panel
75
+ │ ├── billing/ # Stripe billing
76
+ │ ├── tasks/ # Celery tasks
77
+ │ └── utils/ # Helper functions
78
+ ├── tests/ # Test suite (pytest)
79
+ ├── docker-compose.yml # Local development (PostgreSQL, Redis)
80
+ ├── Dockerfile # API container
81
+ ├── Dockerfile.celery # Celery worker container
82
+ ├── pyproject.toml # Dependencies (Poetry) or requirements.txt
83
+ └── .env.example # Environment variables template
84
+ ```
85
+
86
+ ## Current Progress
87
+
88
+ ### Phase 1: Planning & Setup
89
+
90
+ - [x] **Task 1.1**: Database schema design
91
+ - 11 core tables with relationships
92
+ - Comprehensive indexes for performance
93
+ - Row Level Security (RLS) policies
94
+ - Database functions and views
95
+ - Full ER diagram with table documentation
96
+ - [x] **Task 1.2**: FastAPI project structure initialization
97
+ - Core configuration module (`src/core/config.py`) with environment-based settings
98
+ - Logging setup (`src/core/logger.py`) with console and file handlers
99
+ - Main FastAPI application (`src/main.py`) with health check and CORS
100
+ - Dependencies in `requirements.txt`: fastapi, uvicorn, sqlalchemy, psycopg2-binary, pydantic, python-dotenv, bcrypt, python-jose, passlib
101
+ - Environment variable template (`.env.example`)
102
+ - [ ] **Task 1.3**: PostgreSQL setup with SQLAlchemy + Alembic
103
+ - [x] **Task 1.4**: API documentation outline
104
+
105
+ ### Phase 2: Core Backend & Data Model
106
+
107
+ - [x] **Task 2.1**: User authentication system with JWT tokens
108
+ - Implemented `/api/auth/register` endpoint for user signup
109
+ - Implemented `/api/auth/login` endpoint for JWT authentication (access + refresh tokens)
110
+ - Implemented `/api/auth/refresh` endpoint for token rotation
111
+ - Implemented `/api/auth/logout` endpoint to revoke refresh tokens
112
+ - Bcrypt password hashing with passlib
113
+ - JWT token generation with expiration and payload validation
114
+ - Refresh token storage with single-use rotation in database
115
+ - Subscription limit enforcement middleware for free tier users
116
+ - Protected route dependencies (`get_current_user`, `get_current_admin_user`)
117
+ - Comprehensive test suite covering registration, login, token refresh, logout
118
+ - Test fixtures with SQLite in-memory database for isolated tests
119
+ - [x] **Task 2.2**: Invoice management endpoints
120
+ - Full CRUD operations: create, list, update status, soft delete
121
+ - Filtering by status and due date range
122
+ - Pydantic schemas for request/response validation
123
+ - Integration with SQLAlchemy models and database
124
+ - Comprehensive tests with pytest
125
+ - [x] **Task 2.3**: Webhook receivers for Stripe and PayPal
126
+ - Signature verification for both providers
127
+ - Idempotent processing with webhook_events table
128
+ - Automatic invoice status updates
129
+ - Endpoints: `/api/webhooks/stripe` and `/api/webhooks/paypal`
130
+ - Comprehensive logging and error handling
131
+ - Support for key events: invoice.payment_failed, charge.dispute.created, PAYMENT.DENIED, DISPUTE.CREATED
132
+ - [x] **Task 2.4**: Celery background task setup
133
+ - Redis broker and Celery worker configuration
134
+ - Docker Compose setup with Redis service
135
+ - Periodic sync task (`tasks/sync.py`) for invoice status synchronization (fallback if webhooks fail)
136
+ - 15-minute polling interval via Celery Beat
137
+ - Integration with Stripe and PayPal APIs
138
+ - Per-user credential handling for PayPal connections
139
+ - System metrics recording for monitoring
140
+
141
+ ### Phase 3: Integrations & AI Features
142
+
143
+ - [x] **Task 3.1**: Payment provider integration layer
144
+ - Structured clients for Stripe (`integrations/stripe_client.py`), PayPal (`integrations/paypal_client.py`), and Plaid (`integrations/plaid_client.py`)
145
+ - Unified interface: `connect_account()`, `get_invoice_status()`, `list_transactions()`
146
+ - Error handling and retry logic built-in
147
+ - Ready for API credential injection via encrypted database storage
148
+ - [x] **Task 3.2**: Email automation system with follow-up sequences ✓ **COMPLETED**
149
+ - **Email sender abstraction**: `SMTPSender` and `SendGridSender` (placeholder) implementing common interface
150
+ - **Template rendering**: Jinja2-based renderer with HTML and plain text support
151
+ - **Follow-up sequences**: Three automated templates
152
+ - 3-day reminder: Polite pre-due notification
153
+ - 7-day overdue: Clear overdue notice with red warning
154
+ - 14-day final notice: Urgent final warning before dispute escalation
155
+ - **Campaign tracking**: `Campaign` and `EmailEvent` models for send/delivery/engagement metrics
156
+ - **Celery integration**: Asynchronous background tasks (`tasks/email_tasks.py`)
157
+ - `send_followup_task`: Wrapper for single email sending
158
+ - `process_followup_campaign`: Bulk processor for eligible invoices
159
+ - `scheduled_followups`: Periodic task to run every 15 minutes (via Celery Beat)
160
+ - **Database-backed templates**: Template model supports system defaults and user customizations
161
+ - **Comprehensive tests** (`tests/test_email.py`):
162
+ - EmailTemplateRenderer with Jinja2
163
+ - SMTPSender with mocked SMTP
164
+ - `send_followup_email` function with database fixtures
165
+ - Celery task logic and queuing behavior
166
+ - Template file existence and variable validation
167
+ - **Professional email templates**: Responsive HTML templates with inline styling, company branding, payment buttons, and legal compliance footers
168
+ - [x] **Task 3.3**: AI-powered dispute letter drafting
169
+ - **Dispute letter generator** (`src/ai/dispute_drafter.py`) with support for:
170
+ - OpenAI GPT-4 Turbo (`gpt-4-turbo-preview`)
171
+ - Anthropic Claude 3 Opus (`claude-3-opus-20240229`)
172
+ - **Three letter types**:
173
+ - `formal_dispute`: Professional payment dispute letter with 14-day deadline
174
+ - `small_claims_prep`: Court preparation document with case summary and evidence
175
+ - `demand_letter`: Urgent final demand with 7-day deadline and payment plan option
176
+ - **Prompt engineering**: Context-aware prompts that include:
177
+ - Invoice details (number, client, amount, dates, status)
178
+ - Evidence timeline (communications, payments, attempts)
179
+ - Legal references and formatting requirements
180
+ - Output format specification (HTML)
181
+ - **Cost tracking**: Automatic token counting and USD cost calculation
182
+ - GPT-4 Turbo: $0.01/1K input + $0.03/1K output
183
+ - Claude 3 Opus: $0.015/1K input + $0.075/1K output
184
+ - **Provider fallback**: Configurable preference (OpenAI prioritized if both keys present)
185
+ - **Comprehensive test suite** (`tests/test_dispute_drafter.py`):
186
+ - Provider initialization and error handling
187
+ - Prompt building verification for each letter type
188
+ - API mocking for both OpenAI and Anthropic
189
+ - Cost calculation accuracy
190
+ - Evidence formatting and recipient handling
191
+ - [x] **Task 3.4**: PDF generation for dispute letters and small claims forms ✓ **COMPLETED**
192
+ - **PDF Generator** (`src/pdf/generator.py`) using WeasyPrint (HTML → PDF)
193
+ - **Two professional templates**:
194
+ - `dispute_letter.html`: Formal dispute letter with sender/recipient info, invoice details, evidence log, payment instructions, and legal notices
195
+ - `small_claims_form.html`: Court preparation packet with case summary, timeline, evidence list, damages calculation, witness information, filing instructions, and presentation tips
196
+ - **Template rendering**: Jinja2-based with strict variable checking to ensure all required context is provided
197
+ - **Flexible output**: Save to file path or return PDF bytes for S3/cloud storage upload
198
+ - **Customization**: Support for custom CSS, page sizes (A4, Letter), and margins
199
+ - **Comprehensive tests** (`tests/test_pdf.py`) with 11 passing tests:
200
+ - PDF generation from raw HTML
201
+ - Template rendering with context
202
+ - Error handling for missing variables/templates
203
+ - Custom formatting (margins, page size, CSS)
204
+ - Multiple generation scenarios
205
+
206
+
207
+ ### Completed Deliverables
208
+
209
+ 1. **`schema.sql`**: Complete PostgreSQL migration script with:
210
+ - Users, invoices, payment connections, templates
211
+ - Campaigns, A/B tests, email events, webhooks
212
+ - Audit logs, refresh tokens, system metrics
213
+ - Triggers for auto-updating timestamps
214
+ - Functions for business logic (`check_invoice_limit`, `calculate_recovery_rate`, `get_overdue_count`)
215
+ - Materialized views for common queries
216
+ - Encryption extension support for security
217
+
218
+ 2. **`docs/database-er-diagram.md`**: Detailed documentation including:
219
+ - Visual ER diagram with relationship cardinalities
220
+ - Table descriptions with column types and purposes
221
+ - Index strategy and performance notes
222
+ - Security policies (RLS)
223
+ - Migration and rollback instructions
224
+
225
+ ## Getting Started (Local Development)
226
+
227
+ ### Running the Dashboard
228
+
229
+ The Streamlit dashboard provides a user-friendly interface for managing invoices and tracking campaigns.
230
+
231
+ ```bash
232
+ # Install dashboard dependencies
233
+ pip install streamlit streamlit-authenticator plotly pandas
234
+
235
+ # Set API URL (optional, defaults to http://localhost:8000)
236
+ export API_BASE_URL=http://localhost:8000
237
+
238
+ # Run the dashboard
239
+ streamlit run dashboard/app.py
240
+ ```
241
+
242
+ Access the dashboard at http://localhost:8501. Login with credentials you created via the registration API at http://localhost:8000/api/auth/register.
243
+
244
+ **Dashboard Features:**
245
+ - **Overview**: View KPIs (total invoices, recovery rate, overdue count)
246
+ - **Invoices**: Browse and filter invoices, update status inline
247
+ - **Campaigns**: Track email campaign performance and A/B test results
248
+ - **Settings**: Manage payment connections (Stripe, PayPal, Plaid)
249
+
250
+
251
+ ### Prerequisites
252
+
253
+ - Python 3.11+
254
+ - Git
255
+ - Docker & Docker Compose (recommended for database and Redis)
256
+
257
+ ### Quick Start (Using Docker Compose)
258
+
259
+ The fastest way to get started is using Docker Compose, which spins up PostgreSQL and Redis:
260
+
261
+ ```bash
262
+ # 1. Start database and Redis
263
+ docker-compose up -d postgres redis
264
+
265
+ # 2. Install Python dependencies
266
+ pip install -r requirements.txt
267
+
268
+ # 3. Set up environment configuration
269
+ cp .env.example .env
270
+ # Edit .env if needed (defaults work for local Docker setup)
271
+
272
+ # 4. Run database migrations
273
+ # (Alembic will be configured in upcoming tasks; for now, tables are created on startup)
274
+
275
+ # 5. Start the FastAPI server
276
+ uvicorn src.main:app --reload --host 0.0.0.0 --port 8000
277
+
278
+ # 6. In another terminal, start Celery worker
279
+ celery -A src.celery_app worker --loglevel=info
280
+
281
+ # 7. In another terminal, start Celery beat scheduler
282
+ celery -A src.celery_app beat --loglevel=info
283
+ ```
284
+
285
+ Access the API at http://localhost:8000/docs (Swagger UI) or http://localhost:8000/redoc.
286
+
287
+ ### Manual Setup (Without Docker)
288
+
289
+ If you prefer to run PostgreSQL and Redis locally without Docker:
290
+
291
+ ```bash
292
+ # 1. Install and start PostgreSQL and Redis on your system
293
+
294
+ # 2. Create the database
295
+ createdb invoice_resolver
296
+
297
+ # 3. Install Python dependencies
298
+ pip install -r requirements.txt
299
+
300
+ # 4. Configure environment
301
+ cp .env.example .env
302
+ # Edit .env with your local database and Redis URLs
303
+
304
+ # 5. Start services (same as above)
305
+ uvicorn src.main:app --reload
306
+ celery -A src.celery_app worker --loglevel=info
307
+ celery -A src.celery_app beat --loglevel=info
308
+ ```
309
+
310
+ ### Configuration
311
+
312
+ Create a `.env` file from `.env.example` and configure:
313
+
314
+ - **Database**: `DATABASE_URL` (PostgreSQL connection string)
315
+ - **JWT**: `SECRET_KEY` for token signing
316
+ - **Payment APIs**: Stripe, PayPal, Plaid credentials
317
+ - **AI**: `OPENAI_API_KEY` and/or `ANTHROPIC_API_KEY` for dispute letter generation
318
+ - **Email**: SMTP settings for sending follow-ups
319
+ - **Redis**: `REDIS_URL` for Celery task queue
320
+
321
+ See `.env.example` for all available options.
322
+
323
+ ### Project Structure
324
+
325
+ ```
326
+ .
327
+ ├── .env.example # Environment variables template
328
+ ├── requirements.txt # Python dependencies
329
+ ├── src/
330
+ │ ├── main.py # FastAPI application entry point
331
+ │ └── core/
332
+ │ ├── config.py # Settings management (pydantic-settings)
333
+ │ └── logger.py # Logging configuration
334
+ ├── docs/ # Documentation
335
+ ├── schema.sql # Database schema (Phase 1.1)
336
+ └── TASKS.md # Development task list
337
+ ```
338
+
339
+ ### Current Status
340
+
341
+ - ✅ Phase 1.1: Database schema design complete
342
+ - ✅ Phase 1.2: FastAPI project structure initialized
343
+ - 🔄 Phase 1.3: PostgreSQL setup (next task)
344
+ - 🔄 Phase 1.4: API documentation outline (pending)
345
+
346
+ ## API Overview (Planned)
347
+
348
+ ### Public Endpoints
349
+
350
+ - `POST /register` - User registration
351
+ - `POST /login` - JWT authentication
352
+ - `POST /refresh` - Refresh token rotation
353
+
354
+ ### Protected Endpoints
355
+
356
+ - `GET|POST /invoices` - Invoice CRUD with filters
357
+ - `GET /invoices/{id}` - Retrieve invoice details
358
+ - `PATCH /invoices/{id}/status` - Update payment status
359
+ - `POST /templates` - Create custom email templates
360
+ - `GET|POST /campaigns` - Campaign management
361
+ - `GET /analytics/recovery-rate` - KPI metrics
362
+
363
+ ### Webhook Endpoints
364
+
365
+ - `POST /webhooks/stripe` - Stripe event receiver
366
+ - `POST /webhooks/paypal` - PayPal IPN handler
367
+ - `POST /webhooks/plaid` - Plaid transaction sync
368
+
369
+ ### Admin Endpoints (Admin only)
370
+
371
+ - `GET /admin/users` - List all users
372
+ - `GET /admin/metrics` - System health metrics
373
+ - `POST /admin/users/{id}/tier` - Change subscription tier
374
+
375
+ ## Database Schema Highlights
376
+
377
+ ### Tables (11 core tables)
378
+
379
+ | Table | Purpose | Key Columns |
380
+ |-------|---------|-------------|
381
+ | `users` | Account management | subscription_tier, invoice_limit |
382
+ | `payment_connections` | Encrypted credentials | provider (stripe/paypal/plaid) |
383
+ | `invoices` | Invoice records | status, stripe_payment_intent_id |
384
+ | `templates` | Email/document templates | type, variables (JSON) |
385
+ | `campaigns` | Campaign tracking | ab_test_variant, paid_after_send |
386
+ | `ab_tests` | Experiment definitions | test_type, variants (JSON) |
387
+ | `email_events` | Engagement tracking | event_type, occurred_at |
388
+ | `webhook_events` | Idempotency store | provider, event_id (unique) |
389
+ | `audit_logs` | Compliance logging | action, old_values, new_values |
390
+ | `refresh_tokens` | JWT rotation | token_hash, expires_at |
391
+ | `system_metrics` | Time-series data | metric_name, metric_value |
392
+
393
+ ### Key Features
394
+
395
+ - **UUID primary keys** for security and distributed systems
396
+ - **JSONB columns** for flexible metadata storage
397
+ - **Partial indexes** on frequently queried external IDs
398
+ - **Row Level Security** for multi-tenant isolation
399
+ - **Encrypted credentials** for payment provider integrations
400
+ - **Audit trail** for all data changes
401
+ - **Idempotent webhook processing** via unique constraints
402
+ - **Automatic timestamp updates** via triggers
403
+ - **Materialized views** for dashboard performance
404
+
405
+ ## Environment Variables
406
+
407
+ To be documented once implemented (after Task 1.2).
408
+
409
+ ## Testing
410
+
411
+ Test suite will be implemented in Phase 5 using pytest with >80% coverage.
412
+
413
+ ```bash
414
+ pytest tests/ -v --cov=src --cov-report=html
415
+ ```
416
+
417
+ ## Deployment
418
+
419
+ Production deployment instructions will be added in Phase 5 (Docker, cloud hosting, CI/CD).
420
+
421
+ ## License
422
+
423
+ To be determined.
424
+
425
+ ## Contact
426
+
427
+ For questions or feedback about this project, please open an issue on GitHub.
docs/readmes/jira-and-confluence-agent-README.md ADDED
@@ -0,0 +1,306 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # jira-and-confluence-agent
2
+ **Repository:** Julien-ser/jira-and-confluence-agent
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-13T13:12:38Z
8
+
9
+ ---
10
+
11
+ # Jira and Confluence Agent
12
+
13
+ An autonomous agent that creates and modifies Jira and Confluence spaces, similar to what a Jira/Confluence administrator would do manually.
14
+
15
+ ## Features
16
+
17
+ - **Jira Operations**: Create and manage projects, issue types, custom fields, workflows, screens, and permissions
18
+ - **Confluence Operations**: Create and manage spaces, pages, templates, and permissions
19
+ - **Configuration-driven**: Define operations in YAML configuration files
20
+ - **Dry-run mode**: Preview changes before applying them
21
+ - **Idempotent operations**: Safe to run multiple times
22
+ - **Structured logging**: Complete audit trail in `logs/` directory
23
+ - **Detailed reporting**: JSON and text summaries of all operations
24
+
25
+ ## Architecture
26
+
27
+ See [ARCHITECTURE.md](ARCHITECTURE.md) for detailed design and technical specifications.
28
+
29
+ ## Project Structure
30
+
31
+ ```
32
+ jira-and-confluence-agent/
33
+ ├── src/
34
+ │ ├── __init__.py
35
+ │ ├── models.py # Data models and types
36
+ │ ├── parser.py # Configuration parser
37
+ │ ├── jira_client.py # Jira REST API client
38
+ │ ├── confluence_client.py # Confluence REST API client (planned)
39
+ │ ├── engine.py # Operation orchestrator (planned)
40
+ │ ├── reporter.py # Report generation
41
+ │ └── utils.py # Utility functions
42
+ ├── config/ # Configuration files (create your own)
43
+ ├── tests/ # Unit and integration tests
44
+ ├── logs/ # Execution logs and reports
45
+ ├── requirements.txt # Python dependencies
46
+ ├── ARCHITECTURE.md # Architecture documentation
47
+ ├── TASKS.md # Development progress
48
+ └── README.md # This file
49
+ ```
50
+
51
+ ## Setup
52
+
53
+ ### Prerequisites
54
+
55
+ - Python 3.10+
56
+ - Jira/Confluence instance with admin credentials
57
+ - API tokens for authentication
58
+
59
+ ### Installation
60
+
61
+ ```bash
62
+ # Install dependencies using system Python (no virtualenv needed)
63
+ pip install -r requirements.txt
64
+ ```
65
+
66
+ ## Configuration
67
+
68
+ Create a configuration file (YAML format) defining the operations you want to perform:
69
+
70
+ ```yaml
71
+ # example_config.yaml
72
+ connections:
73
+ jira:
74
+ url: "https://your-jira-instance.atlassian.net"
75
+ username: "admin@example.com"
76
+ password: "YOUR_API_TOKEN"
77
+ confluence:
78
+ url: "https://your-confluence-instance.atlassian.net"
79
+ username: "admin@example.com"
80
+ password: "YOUR_API_TOKEN"
81
+
82
+ jira:
83
+ projects:
84
+ - key: "TEST"
85
+ name: "Test Project"
86
+ description: "A test project"
87
+ projectTypeKey: "business"
88
+
89
+ confluence:
90
+ spaces:
91
+ - key: "TEST"
92
+ name: "Test Space"
93
+ description: "A test space"
94
+ ```
95
+
96
+ ## Usage
97
+
98
+ ### Command Line
99
+
100
+ ```bash
101
+ # Run with dry-run to preview changes
102
+ python run_agent.py config/example_config.yaml --dry-run
103
+
104
+ # Execute actual operations
105
+ python run_agent.py config/example_config.yaml
106
+
107
+ # Enable verbose logging
108
+ python run_agent.py config/example_config.yaml --verbose
109
+
110
+ # Custom log directory
111
+ python run_agent.py config/example_config.yaml --log-dir mylogs
112
+ ```
113
+
114
+ ### Programmatic Usage
115
+
116
+ ```python
117
+ from src.parser import OperationParser
118
+ from src.engine import Engine
119
+ from src.reporter import Reporter
120
+
121
+ # Parse configuration
122
+ parser = OperationParser()
123
+ operations = parser.parse_file("config/example.yaml")
124
+
125
+ # Execute operations
126
+ engine = Engine(config["connections"], dry_run=False)
127
+ with engine:
128
+ report = engine.run(operations)
129
+
130
+ # Report is automatically generated in logs/ directory
131
+ ```
132
+
133
+ ## Deployment
134
+
135
+ The Jira and Confluence Agent can be deployed in multiple ways:
136
+
137
+ ### Option 1: Docker (Recommended)
138
+
139
+ The easiest way to deploy is using Docker:
140
+
141
+ ```bash
142
+ # Build the Docker image
143
+ docker build -t jira-confluence-agent .
144
+
145
+ # Run with a specific config file
146
+ docker run -v $(pwd)/config:/app/config:ro \
147
+ -v $(pwd)/logs:/app/logs \
148
+ -e AGENT_CONFIG=/app/config/example_config.yaml \
149
+ jira-confluence-agent
150
+
151
+ # Or use docker-compose for easier setup
152
+ docker-compose up
153
+ ```
154
+
155
+ **Using environment variables with Docker Compose:**
156
+
157
+ 1. Copy `.env.example` to `.env` and fill in your credentials
158
+ 2. Update `config/agent_config.yaml` or use the example config
159
+ 3. Run: `docker-compose up`
160
+
161
+ The container is designed for manual execution only (restart: "no"). For scheduled/automated runs, consider:
162
+
163
+ - Using cron to execute `docker-compose run jira-agent`
164
+ - Setting up a CI/CD pipeline (see below)
165
+
166
+ ### Option 2: Direct Python Execution
167
+
168
+ For simple deployments, run directly on the host:
169
+
170
+ ```bash
171
+ # Install dependencies system-wide
172
+ pip install -r requirements.txt
173
+
174
+ # Run the agent
175
+ python run_agent.py config/example_config.yaml --dry-run
176
+
177
+ # With environment variables (recommended for credentials)
178
+ export JIRA_URL="https://your-company.atlassian.net"
179
+ export JIRA_USERNAME="admin@example.com"
180
+ export JIRA_PASSWORD="your_api_token"
181
+ # ... set other variables as needed
182
+ python run_agent.py config/example_config.yaml
183
+ ```
184
+
185
+ ### Option 3: Package Installation
186
+
187
+ Create a pip-installable package (requires `setup.py` - to be added):
188
+
189
+ ```bash
190
+ pip install -e .
191
+ jira-agent config/example_config.yaml
192
+ ```
193
+
194
+ ### Production Considerations
195
+
196
+ #### Security
197
+ - **Never store credentials in config files** for production. Use environment variables or a secrets manager
198
+ - The agent supports reading credentials from environment variables:
199
+ - `JIRA_URL`, `JIRA_USERNAME`, `JIRA_PASSWORD`
200
+ - `CONFLUENCE_URL`, `CONFLUENCE_USERNAME`, `CONFLUENCE_PASSWORD`
201
+ - Rotate API tokens regularly
202
+ - Use SSL verification (default: true)
203
+
204
+ #### Scheduling Automated Runs
205
+ Use cron (with Docker) or systemd timers:
206
+
207
+ ```bash
208
+ # Example cron job (runs daily at 2 AM)
209
+ 0 2 * * * cd /path/to/agent && docker-compose run --rm jira-agent /app/config/daily_ops.yaml >> /var/log/jira-agent.log 2>&1
210
+ ```
211
+
212
+ #### Monitoring
213
+ - Check `logs/` directory for `agent.log`, `audit.log`, and reports
214
+ - Set up log rotation for `logs/` directory
215
+ - Monitor exit codes: `0` = success, `1` = failure
216
+
217
+ #### High Availability
218
+ - The agent is idempotent and can be safely re-run
219
+ - Design configurations to be modular and reusable
220
+ - Use dry-run mode (`--dry-run`) to validate before production runs
221
+
222
+ ### CI/CD Integration
223
+
224
+ Example GitHub Actions workflow (`.github/workflows/deploy.yml`):
225
+
226
+ ```yaml
227
+ name: Deploy Jira Agent
228
+ on:
229
+ schedule:
230
+ - cron: '0 2 * * *' # Daily at 2 AM
231
+ workflow_dispatch:
232
+
233
+ jobs:
234
+ run-agent:
235
+ runs-on: ubuntu-latest
236
+ steps:
237
+ - uses: actions/checkout@v3
238
+ - name: Set up Python
239
+ uses: actions/setup-python@v4
240
+ with:
241
+ python-version: '3.11'
242
+ - name: Install dependencies
243
+ run: pip install -r requirements.txt
244
+ - name: Run agent
245
+ env:
246
+ JIRA_URL: ${{ secrets.JIRA_URL }}
247
+ JIRA_USERNAME: ${{ secrets.JIRA_USERNAME }}
248
+ JIRA_PASSWORD: ${{ secrets.JIRA_PASSWORD }}
249
+ CONFLUENCE_URL: ${{ secrets.CONFLUENCE_URL }}
250
+ CONFLUENCE_USERNAME: ${{ secrets.CONFLUENCE_USERNAME }}
251
+ CONFLUENCE_PASSWORD: ${{ secrets.CONFLUENCE_PASSWORD }}
252
+ run: |
253
+ python run_agent.py config/production.yaml --dry-run
254
+ python run_agent.py config/production.yaml
255
+ ```
256
+
257
+ ## Current Status
258
+
259
+ **Project Status: Production Ready** ✅
260
+
261
+ All core features implemented and tested:
262
+
263
+ - ✅ Configuration parser (YAML/JSON)
264
+ - ✅ Jira REST API client (projects, issue types, custom fields, workflows)
265
+ - ✅ Confluence REST API client (spaces, pages, templates)
266
+ - ✅ Operation engine with dry-run support
267
+ - ✅ Comprehensive reporting and audit logging
268
+ - ✅ 40 passing unit and integration tests
269
+ - ✅ Command-line interface (run_agent.py)
270
+ - ✅ Example configuration file
271
+
272
+ See [TASKS.md](TASKS.md) for detailed development progress.
273
+
274
+ ## Development
275
+
276
+ Run tests (when available):
277
+ ```bash
278
+ pytest tests/
279
+ ```
280
+
281
+ Check linting (when configured):
282
+ ```bash
283
+ ruff check src/
284
+ ```
285
+
286
+ ## Logging
287
+
288
+ All operations are logged to the `logs/` directory:
289
+ - `agent.log` - Structured application logs
290
+ - `audit.log` - JSON audit trail of each operation
291
+ - `report_<timestamp>.json` - Machine-readable execution report
292
+ - `summary_<timestamp>.txt` - Human-readable summary
293
+
294
+ ## Security
295
+
296
+ - Never commit credentials or API tokens to version control
297
+ - Use environment variables or secure credential stores
298
+ - The `.gitignore` excludes sensitive files
299
+
300
+ ## Contributing
301
+
302
+ This is an autonomous project managed by OpenCode. See TASKS.md for current development status.
303
+
304
+ ## License
305
+
306
+ [Add your license here]
docs/readmes/julien-rag-README.md ADDED
@@ -0,0 +1,298 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # julien-rag
2
+ **Repository:** Julien-ser/julien-rag
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-13T02:30:35Z
8
+
9
+ ---
10
+
11
+ # julien-rag
12
+
13
+ **Mission:** Create a vector database of everything I've done online/GitHub that can be used elsewhere as a RAG implementation.
14
+
15
+ ## Overview
16
+
17
+ This project builds a Retrieval-Augmented Generation (RAG) system that:
18
+ - Collects data from GitHub (repos, commits, issues, PRs, gists, starred)
19
+ - Scrapes web presence (blog, forums, social media)
20
+ - Processes and chunks documents intelligently
21
+ - Stores embeddings in a vector database
22
+ - Provides a FastAPI REST interface for semantic search and Q&A
23
+ - Can be used as a Python SDK in other projects
24
+
25
+ ## Technology Stack
26
+
27
+ - **Vector Database**: ChromaDB (local, persistent)
28
+ - **Embeddings**: OpenAI `text-embedding-ada-002` or `sentence-transformers/all-MiniLM-L6-v2`
29
+ - **API**: FastAPI with async endpoints
30
+ - **Data Collection**: PyGithub, beautifulsoup4, requests
31
+ - **Processing**: tiktoken for token counting, recursive text splitting
32
+
33
+ ## Current Status
34
+
35
+ **Phase 1: Planning & Infrastructure Setup** ✅ Complete
36
+ - [x] **Task 1.1**: Vector database selection (ChromaDB chosen for local-first, zero-config approach)
37
+ - [x] **Task 1.2**: Design data schema and document structure
38
+ - [x] **Task 1.3**: Choose embedding model and API setup
39
+ - [x] **Task 1.4**: Initialize project structure and dependencies
40
+
41
+ **Phase 2: Data Collection & Ingestion Pipeline** ✅ Complete
42
+ - [x] **Task 2.1**: Implement GitHub API data collector ✅
43
+ - [x] **Task 2.2**: Implement web content scraper for online presence ✅
44
+ - [x] **Task 2.3**: Build document preprocessing and chunking pipeline ✅
45
+ - [x] **Task 2.4**: Create unified data pipeline with error handling ✅
46
+ - `src/pipeline.py` with comprehensive orchestration
47
+ - Retry logic, incremental updates, detailed logging
48
+ - Shell script: `scripts/ingest_all.sh`
49
+ - Logs written to `logs/ingestion_*.log`
50
+ - Statistics saved to `data/processed/pipeline_stats.json`
51
+
52
+ **Phase 3: Vector Database Implementation** ✅ Complete
53
+ - [x] **Task 3.1**: Initialize vector database and collections ✅
54
+ - [x] **Task 3.2**: Implement embedding generation and storage ✅
55
+ - [x] **Task 3.3**: Implement similarity search functionality ✅
56
+ - [x] **Task 3.4**: Perform database validation and optimization ✅
57
+ - `scripts/validate_db.py` with comprehensive validation suite
58
+ - `docs/database_performance.md` with performance metrics and recommendations
59
+ - Tests: data integrity, latency benchmarks, metadata filtering, recall@k support
60
+
61
+ **Phase 4: RAG API & External Integration** ✅ Complete
62
+ - [x] **Task 4.1**: Build FastAPI REST endpoints ✅
63
+ - Complete API with interactive docs at `/docs`
64
+ - All endpoints: `/query`, `/sources`, `/stats`, `/refresh`, `/health`, `/metrics`, `/collections`
65
+ - Async support, CORS enabled, admin authentication, comprehensive error handling
66
+ - 22/22 API unit tests passing
67
+
68
+ - [x] **Task 4.2**: Implement RAG generation pipeline ✅
69
+ - `src/rag.py` with `RAGPipeline` class supporting OpenAI and local providers
70
+ - API endpoint `/rag-query` returns `{answer, confidence, sources, stats}`
71
+ - Configuration in `config/rag.yaml` with LLM settings and prompts
72
+ - 22/23 tests passing (1 requires optional dependency)
73
+
74
+ - [x] **Task 4.3**: Create SDK/client library for external use ✅
75
+ - Complete Python package `julien_rag` with `RAGClient` class
76
+ - Supports all API endpoints with typed Pydantic models
77
+ - Full test suite (15/15 passing)
78
+ - Usage examples in `examples/usage_example.py`
79
+ - Documentation in README and package
80
+
81
+ - [x] **Task 4.4**: Add monitoring, logging, and deployment configuration ✅
82
+ - `src/monitoring.py` with Prometheus metrics (Counter, Histogram, Gauge) and `/metrics` endpoint
83
+ - Metrics middleware for automatic request tracking (latency, status codes, endpoints)
84
+ - Database query metrics (operation duration, document counts)
85
+ - Embedding generation metrics (token count, duration, provider)
86
+ - RAG pipeline metrics (query count, confidence scores)
87
+ - Optional system metrics (CPU, memory) when psutil available
88
+ - Production-ready `docker/Dockerfile` with multi-stage build, health checks, non-root user
89
+ - `docker/docker-compose.yml` with API service and optional monitoring stack (Prometheus, Grafana)
90
+ - Comprehensive integration test suite in `tests/integration/test_full_flow.py`
91
+ - Deployment guide in `docs/deployment.md`
92
+
93
+ **MISSION ACCOMPLISHED** ✅
94
+ Vector DB with full RAG implementation ready for external use.
95
+
96
+ ## Getting Started
97
+
98
+ ### Prerequisites
99
+
100
+ - Python 3.9+
101
+ - Git
102
+ - (Optional) GitHub API token for data collection: `GITHUB_TOKEN`
103
+ - (Optional) OpenAI API key for embeddings: `OPENAI_API_KEY`
104
+ - (Optional) For Docker deployment: Docker and Docker Compose
105
+
106
+ ### Option 1: Local Python Installation
107
+
108
+ ```bash
109
+ # Clone and navigate
110
+ cd projects/julien-rag
111
+
112
+ # Install dependencies
113
+ pip install -r requirements.txt
114
+
115
+ # Set up environment variables
116
+ cp .env.example .env # if .env.example exists
117
+ # Edit .env with your API keys:
118
+ # - GITHUB_TOKEN (optional, for GitHub data collection)
119
+ # - OPENAI_API_KEY (optional, for embeddings and RAG with OpenAI)
120
+ # - ADMIN_TOKEN (optional, for protected refresh endpoint)
121
+
122
+ # Run data ingestion (collect and process data)
123
+ ./scripts/ingest_all.sh
124
+ # Or: python -m src.pipeline
125
+
126
+ # Start the FastAPI server
127
+ uvicorn src.api:app --reload --port 8000
128
+
129
+ # Visit http://localhost:8000/docs for interactive API documentation
130
+ ```
131
+
132
+ ### Option 2: Docker Deployment (Recommended for Production)
133
+
134
+ ```bash
135
+ # Build and run using Docker Compose
136
+ docker-compose up -d
137
+
138
+ # Or build manually:
139
+ docker build -t julien-rag -f docker/Dockerfile .
140
+ docker run -p 8000:8000 \
141
+ -v $(pwd)/data:/home/app/data \
142
+ -v $(pwd)/logs:/home/app/logs \
143
+ -v $(pwd)/config:/home/app/config:ro \
144
+ -e OPENAI_API_KEY=${OPENAI_API_KEY} \
145
+ -e ADMIN_TOKEN=${ADMIN_TOKEN} \
146
+ julien-rag
147
+
148
+ # Access API at http://localhost:8000/docs
149
+ ```
150
+
151
+ **Optional Monitoring Stack:**
152
+
153
+ ```bash
154
+ # Deploy with Prometheus and Grafana for metrics
155
+ docker-compose --profile monitoring up -d
156
+
157
+ # Access services:
158
+ # - API: http://localhost:8000
159
+ # - Prometheus: http://localhost:9090
160
+ # - Grafana: http://localhost:3000 (admin/admin)
161
+ ```
162
+
163
+ For detailed deployment options, see [docs/deployment.md](docs/deployment.md).
164
+
165
+ ### Running the Ingestion Pipeline
166
+
167
+ The unified pipeline orchestrates data collection, preprocessing, and chunk generation.
168
+
169
+ **Using the shell script (recommended):**
170
+
171
+ ```bash
172
+ # Full ingestion with default settings
173
+ ./scripts/ingest_all.sh
174
+
175
+ # With custom configuration
176
+ ./scripts/ingest_all.sh --config config/pipeline_config.json
177
+
178
+ # Set log level (DEBUG, INFO, WARNING, ERROR)
179
+ ./scripts/ingest_all.sh --log-level DEBUG
180
+ ```
181
+
182
+ **Using Python module directly:**
183
+
184
+ ```bash
185
+ # Run with default configuration
186
+ python -m src.pipeline
187
+
188
+ # With configuration file
189
+ python -m src.pipeline --config config/my_config.json
190
+
191
+ # With debug logging
192
+ python -m src.pipeline --log-level DEBUG
193
+ ```
194
+
195
+ **Configuration:**
196
+
197
+ Create a JSON configuration file (see `config/pipeline_config.example.json`):
198
+
199
+ ```json
200
+ {
201
+ "incremental": true,
202
+ "chunk_size": 512,
203
+ "chunk_overlap": 100,
204
+ "github_token": null,
205
+ "web_scrape_config": {
206
+ "personal": ["https://yourwebsite.com/about"],
207
+ "blog": ["https://yourblog.com/rss"]
208
+ }
209
+ }
210
+ ```
211
+
212
+ **Logs & Statistics:**
213
+
214
+ - Logs: `logs/ingestion_YYYYMMDD_HHMMSS.log` (rotating, max 10MB per file)
215
+ - Statistics: `data/processed/pipeline_stats.json`
216
+ - Processed chunks: `data/processed/*_chunks.jsonl`
217
+
218
+ **Features:**
219
+
220
+ - ✅ Retry logic with exponential backoff for API calls
221
+ - ✅ Incremental updates (skips unchanged files)
222
+ - ✅ Graceful error handling - continues on failures
223
+ - ✅ Comprehensive logging with rotation
224
+ - ✅ Automatic fallback to existing raw data
225
+ - ✅ Statistics and performance metrics
226
+
227
+ ## Project Structure
228
+
229
+ ```
230
+ julien-rag/
231
+ ├── src/
232
+ │ ├── database.py # ✅ ChromaDB initialization (Task 3.1)
233
+ │ ├── embedder.py # ✅ Embedding generation (Task 3.2)
234
+ │ ├── vector_store.py # ✅ Vector storage operations (Task 3.2)
235
+ │ ├── retriever.py # Similarity search (Task 3.3 - pending)
236
+ │ ├── github_collector.py
237
+ │ ├── web_scraper.py
238
+ │ ├── preprocessor.py
239
+ │ ├── pipeline.py
240
+ │ ├── api.py # FastAPI endpoints (Task 4.1 - pending)
241
+ │ ├── rag.py # RAG generation (Task 4.2 - pending)
242
+ │ └── monitoring.py # Monitoring (Task 4.4 - pending)
243
+ ├── data/
244
+ │ ├── raw/ # Raw collected data
245
+ │ ├── processed/ # Chunked documents
246
+ │ └── vector_db/ # ChromaDB storage
247
+ ├── config/
248
+ │ ├── embeddings.yaml # Embedding configuration
249
+ │ └── rag.yaml # RAG configuration (pending)
250
+ ├── tests/
251
+ │ ├── test_embedder.py # ✅ All tests passing (54/54)
252
+ │ ├── test_vector_store.py # ✅ Vector store tests
253
+ │ ├── test_database.py # ✅ Database tests
254
+ │ ├── test_preprocessor.py # ✅ Preprocessor tests
255
+ │ ├── test_web_scraper.py # ✅ Web scraper tests
256
+ │ └── test_github_collector.py # ✅ GitHub collector tests
257
+ ├── docs/
258
+ │ ├── vector_db_selection.md # ✅ Completed
259
+ │ ├── schema_design.md # ✅ Completed
260
+ │ ├── database_performance.md # ✅ Completed (Task 3.4)
261
+ │ └── deployment.md # Pending (Task 4.4)
262
+ ├── logs/
263
+ ├── scripts/
264
+ │ ├── ingest_all.sh # Full ingestion pipeline
265
+ │ └── validate_db.py # Database validation & benchmarking
266
+ ├── examples/
267
+ │ └── github_collector_example.py
268
+ ├── requirements.txt
269
+ ├── TASKS.md
270
+ └── README.md
271
+ ```
272
+
273
+ ## Using the SDK
274
+
275
+ The SDK is now available as a Python package! Install it and use the RAG API in any project.
276
+
277
+ ### SDK Installation
278
+
279
+ ```bash
280
+ # Install the SDK from the local package
281
+ pip install -e .
282
+
283
+ # Or install from a published package (once available)
284
+ # pip install julien_rag
285
+ ```
286
+
287
+ ### SDK Usage
288
+
289
+ See [docs/deployment.md](docs/deployment.md) for detailed deployment options including:
290
+ - Docker deployment with docker-compose
291
+ - Production configuration
292
+ - Monitoring with Prometheus/Grafana
293
+ - Security best practices
294
+ - Troubleshooting
295
+
296
+ ---
297
+
298
+ ## Project Structure
docs/readmes/mlflow-ai-experiment-README.md ADDED
@@ -0,0 +1,379 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # mlflow-ai-experiment
2
+ **Repository:** Julien-ser/mlflow-ai-experiment
3
+ **Contribution:** 👑 Owner
4
+ **Stars:** 0
5
+ **Description:** None
6
+ **Language:** Python
7
+ **Last Updated:** 2026-03-19T12:51:49Z
8
+
9
+ ---
10
+
11
+ # MLFlow AI Experiment: Text Classification Model Comparison
12
+
13
+ ## Overview
14
+ This project uses **MLFlow** to systematically compare state-of-the-art machine learning models for sentiment analysis on the IMDB movie reviews dataset. We evaluate both traditional ML approaches and modern transformer-based architectures to identify the optimal balance of accuracy, speed, and resource efficiency.
15
+
16
+ ## Problem Statement
17
+ **Task**: Binary text classification (positive/negative sentiment) on 50,000 movie reviews.
18
+
19
+ **Goal**: Compare multiple model architectures and determine the best-performing model(s) based on:
20
+ - Accuracy and F1 score
21
+ - Inference latency
22
+ - Model size and memory footprint
23
+ - Training efficiency
24
+
25
+ **Detailed problem statement**: [docs/problem-statement.md](docs/problem-statement.md)
26
+
27
+ ## Project Structure
28
+ ```
29
+ .
30
+ ├── README.md # Project documentation
31
+ ├── TASKS.md # Development task tracking
32
+ ├── requirements.txt # Python dependencies
33
+ ├── config.yaml # MLflow and experiment configuration
34
+ ├── setup_mlflow.py # MLflow tracking setup script
35
+ ├── .github/workflows/ # CI/CD pipelines
36
+ │ └── test.yml
37
+ ├── docs/ # Documentation and problem statement
38
+ │ └── problem-statement.md
39
+ ├── src/ # Source code
40
+ │ ├── models/ # Model implementations
41
+ │ ├── training.py # Training pipeline
42
+ │ ├── evaluation.py # Metrics computation
43
+ │ └── ...
44
+ ├── data/ # Processed datasets
45
+ ├── models/ # Saved model artifacts
46
+ ├── experiments/ # Experiment configurations
47
+ ├── notebooks/ # Jupyter notebooks for exploration
48
+ └── mlruns/ # MLFlow tracking data (auto-generated)
49
+ ```
50
+
51
+ ## Setup
52
+
53
+ ### Prerequisites
54
+ - Python 3.9+
55
+ - Git
56
+
57
+ ### Option 1: Using pip (system Python)
58
+ ```bash
59
+ # Install dependencies directly into system Python
60
+ pip install -r requirements.txt
61
+
62
+ # Verify environment
63
+ python verify_environment.py
64
+ ```
65
+
66
+ ### Option 2: Using Conda (recommended for isolation)
67
+ ```bash
68
+ # Create and activate conda environment
69
+ conda env create -f environment.yml
70
+ conda activate mlflow-ai-experiment
71
+
72
+ # Verify environment
73
+ python verify_environment.py
74
+ ```
75
+
76
+ ### Quick Start
77
+ 1. Start MLflow tracking server:
78
+ ```bash
79
+ mlflow ui
80
+ ```
81
+ Open http://localhost:5000 in your browser.
82
+
83
+ 2. Explore the notebooks:
84
+ ```bash
85
+ jupyter lab notebooks/01_data_exploration.ipynb
86
+ ```
87
+
88
+ 3. Run baseline training:
89
+ ```bash
90
+ python src/train.py --model logistic_regression --config config/baseline.yaml
91
+ ```
92
+
93
+ ## Usage
94
+
95
+ ### Starting MLFlow UI
96
+ ```bash
97
+ mlflow ui
98
+ ```
99
+ Then open http://localhost:5000 in your browser.
100
+
101
+ ### Running Experiments
102
+
103
+ #### Baseline Model (TF-IDF + Logistic Regression)
104
+ ```bash
105
+ python scripts/run_baseline.py
106
+ ```
107
+
108
+ #### Classical ML Models
109
+ ```bash
110
+ # Train baseline model using classical pipeline
111
+ python src/train.py --model logistic_regression --config config.yaml
112
+
113
+ # Train other classical models
114
+ python src/train.py --model svm --config config.yaml
115
+ python src/train.py --model random_forest --config config.yaml
116
+ python src/train.py --model xgboost --config config.yaml
117
+ ```
118
+
119
+ Or train all classical models at once with the comprehensive script:
120
+ ```bash
121
+ python scripts/run_classical_models.py
122
+ ```
123
+
124
+ This script trains all classical models (Logistic Regression, SVM, Random Forest, XGBoost) and generates a comparison table with all metrics logged to MLflow.
125
+
126
+ #### Transformer Models
127
+ The project now includes a unified transformer interface supporting multiple architectures:
128
+
129
+ ```bash
130
+ # BERT base or large
131
+ python src/train.py --model bert --config config/transformer_bert.yaml
132
+ python src/train.py --model bert-large --config config/transformer_bert.yaml
133
+
134
+ # RoBERTa
135
+ python src/train.py --model roberta --config config/transformer_roberta.yaml
136
+
137
+ # DeBERTa (v3)
138
+ python src/train.py --model deberta --config config/transformer_deberta.yaml
139
+
140
+ # XLNet
141
+ python src/train.py --model xlnet --config config/transformer_xlnet.yaml
142
+
143
+ # Or use custom HuggingFace model names
144
+ python src/train.py --model distilbert-base-uncased --config config/transformer_distilbert.yaml
145
+ python src/train.py --model google/electra-base-discriminator --config config/transformer_electra.yaml
146
+ ```
147
+
148
+ The transformer models are implemented in `src/models/transformers.py` with:
149
+ - Unified `TransformerModel` base class with consistent API
150
+ - Specific wrappers: `BERTModel`, `RoBERTaModel`, `DeBERTaModel`, `XLNetModel`
151
+ - Factory functions: `create_transformer_model()` and `create_transformer_model_from_name()`
152
+ - Support for custom classification heads with configurable dropout
153
+ - Automatic MLflow logging with transformers flavor
154
+
155
+ ### Viewing Dashboards & Analysis
156
+
157
+ #### Interactive Streamlit Dashboard
158
+ An interactive dashboard for model comparison and analysis:
159
+ ```bash
160
+ streamlit run app/dashboard.py
161
+ ```
162
+ Then open http://localhost:8501 in your browser.
163
+
164
+ The dashboard provides:
165
+ - Performance comparison across all models (bar charts)
166
+ - Latency vs accuracy trade-off analysis
167
+ - Model size vs performance visualizations
168
+ - Metric correlation heatmaps
169
+ - Statistical significance testing (Friedman test, bootstrap CIs)
170
+ - Exportable results tables
171
+
172
+ **Note**: The dashboard automatically loads data from MLflow. If insufficient data exists, it will display sample data for demonstration.
173
+
174
+ #### Jupyter Notebook Analysis
175
+ For detailed exploratory analysis:
176
+ ```bash
177
+ jupyter lab notebooks/analysis.ipynb
178
+ ```
179
+
180
+ The analysis notebook includes:
181
+ - Comprehensive data exploration
182
+ - Statistical testing (Friedman, Nemenyi, bootstrap)
183
+ - Correlation analysis
184
+ - Production deployment recommendations
185
+ - Key insights and visualizations
186
+
187
+ #### Documentation
188
+ Detailed model comparison report: [docs/model_comparison.md](docs/model_comparison.md)
189
+
190
+ ### Evaluating Models
191
+
192
+ The project includes a comprehensive automated evaluation suite that computes all metrics (accuracy, precision, recall, F1, confusion matrix, inference latency, memory footprint) and logs them consistently.
193
+
194
+ #### Single Model Evaluation
195
+ ```bash
196
+ python src/evaluate.py --model-path models/best_model/
197
+ ```
198
+
199
+ #### Batch Evaluation & Comparison
200
+ Evaluate all models from an MLflow experiment and generate comparison tables:
201
+ ```bash
202
+ python scripts/evaluate_all.py \
203
+ --experiment-name "my_experiment" \
204
+ --test-data data/test.csv \
205
+ --output-dir comparison_results
206
+ ```
207
+
208
+ Or evaluate specific runs:
209
+ ```bash
210
+ python scripts/evaluate_all.py \
211
+ --run-ids <run_id1> <run_id2> \
212
+ --test-data data/test.csv
213
+ ```
214
+
215
+ All results are automatically logged to MLflow and saved as CSV/JSON comparison tables.
216
+
217
+ ## Data Pipeline Performance Benchmark
218
+
219
+ The project includes a comprehensive benchmark to measure data loading and preprocessing performance across different batch sizes and configurations.
220
+
221
+ ### Running the Benchmark
222
+
223
+ ```bash
224
+ # Install psutil if not already installed
225
+ pip install psutil
226
+
227
+ # Run the benchmark (logs results to MLFlow)
228
+ python scripts/benchmark_data.py
229
+ ```
230
+
231
+ The benchmark measures:
232
+ - **Data loading**: Time and memory for loading IMDB dataset with different batch sizes
233
+ - **Classical preprocessing**: TF-IDF vectorization with various feature counts
234
+ - **Transformer tokenization**: BERT, RoBERTa, DistilBERT tokenization throughput
235
+
236
+ All results are automatically logged to MLFlow in the `data_pipeline_benchmarks` experiment.
237
+
238
+ ### Viewing Results
239
+
240
+ ```bash
241
+ # Start MLFlow UI
242
+ mlflow ui
243
+ # Open http://localhost:5000
244
+ ```
245
+
246
+ Detailed analysis and recommendations are available in:
247
+ - [docs/data_performance.md](docs/data_performance.md)
248
+
249
+ ### Benchmark Outputs
250
+
251
+ The script generates CSV files with detailed metrics and uploads them as MLFlow artifacts:
252
+ - `data_loading_benchmark.csv`
253
+ - `classical_preprocessing_benchmark.csv`
254
+ - `transformer_preprocessing_benchmark.csv`
255
+
256
+ ## Model Categories
257
+
258
+ ### Classical ML Models
259
+ - Logistic Regression (baseline)
260
+ - Support Vector Machines (SVM)
261
+ - Random Forest
262
+ - XGBoost
263
+ - LightGBM
264
+
265
+ ### Transformer Models
266
+ The project includes a **unified transformer interface** with support for:
267
+
268
+ **Core architectures** (directly implemented):
269
+ - BERT (base, large, and any HuggingFace variant)
270
+ - RoBERTa (base, large)
271
+ - DeBERTa (v3 base/large)
272
+ - XLNet (base, large)
273
+
274
+ **Extended support** (any HuggingFace model):
275
+ - DistilBERT
276
+ - ELECTRA
277
+ - ALBERT
278
+ - GPT-2/3 for classification
279
+ - And any other text-classification model from HuggingFace Hub
280
+
281
+ All transformer models share the same API via `TransformerModel` base class with:
282
+ - Consistent training, prediction, and logging interfaces
283
+ - Customizable classification heads
284
+ - Automatic device management (CPU/GPU)
285
+ - MLflow transformers flavor integration
286
+
287
+ ## Tracking Experiments
288
+ All experiments are automatically tracked in MLFlow with:
289
+ - Model parameters
290
+ - Performance metrics
291
+ - System metrics (inference time, memory usage)
292
+ - Model artifacts
293
+ - Dataset version information
294
+
295
+ ## Data Utilities & Versioning
296
+
297
+ ### Data Versioning (`src/data_versioning.py`)
298
+ The project implements checksum-based data versioning to ensure reproducibility:
299
+ - **Automatic version detection**: SHA-256 hashes of dataset content
300
+ - **Version manifest**: `data/version_manifest.yaml` stores checksums and timestamps
301
+ - **Integrity verification**: `verify_dataset_integrity()` ensures datasets haven't changed
302
+ - **Version string**: Compact format (e.g., `v1.0-a1b2c3d4`) for tracking
303
+
304
+ ### MLFlow Data Logging (`src/data_utils.py`)
305
+ Comprehensive utilities to log data information to MLFlow:
306
+ - `log_dataset_statistics()`: Sample counts, class distribution, text lengths
307
+ - `log_preprocessing_parameters()`: Track preprocessing configuration
308
+ - `log_data_artifacts()`: Save dataset splits as artifacts (CSV/Parquet)
309
+ - `create_and_log_data_report()`: Generate markdown data reports
310
+ - `prepare_data_for_mlflow()`: All-in-one function for complete data logging
311
+
312
+ ### Usage Example
313
+ ```python
314
+ from mlflow_ai_experiment.data_loader import load_and_log_dataset
315
+
316
+ # Load data and automatically log to MLFlow with versioning
317
+ train_df, val_df, test_df = load_and_log_dataset(
318
+ log_to_mlflow=True,
319
+ preprocessing_params={
320
+ "cleaning": "lowercase, remove_html",
321
+ "tokenization": "bert-base-uncased",
322
+ }
323
+ )
324
+ ```
325
+
326
+ This automatically:
327
+ - Calculates dataset checksums
328
+ - Creates `data/version_manifest.yaml`
329
+ - Logs statistics, parameters, and artifacts to MLFlow
330
+ - Generates a comprehensive data report
331
+
332
+ ## Current Status
333
+
334
+ **Phase 1: Planning & Setup** - ✓ Complete
335
+ - [x] Problem statement and requirements defined (see [docs/problem-statement.md](docs/problem-statement.md))
336
+ - [x] MLFlow tracking infrastructure setup
337
+ - [x] Development environment creation
338
+ - [x] Baseline model implementation
339
+
340
+ **Phase 2: Data Management & Preprocessing** - ✓ Complete
341
+ - [x] Dataset download and preparation
342
+ - [x] Text preprocessing pipeline
343
+ - [x] Data utilities for MLFlow logging
344
+ - [x] Data pipeline performance benchmarking
345
+
346
+ **Phase 3: Model Implementation & Training** - ✓ Complete
347
+ - [x] HuggingFace transformer integration
348
+ - [x] Classical ML model implementations
349
+ - [x] State-of-the-art models (ELECTRA, ALBERT, DistilBERT)
350
+ - [x] Unified training pipeline
351
+
352
+ **Phase 4: Experimentation, Logging & Analysis** - ✓ Complete
353
+ - [x] Comprehensive MLFlow experiment tracking
354
+ - [x] Hyperparameter optimization framework
355
+ - [x] Automated model evaluation suite
356
+ - [x] Interactive dashboards and visualizations
357
+ - [x] Final report and recommendations
358
+ - [x] Reproducibility framework
359
+ - [x] Model artifact management and export utilities
360
+
361
+ **Project Status: 100% COMPLETE** ✅
362
+
363
+ All deliverables have been successfully implemented and documented. See the [Final Report](docs/FINAL_REPORT.md) for comprehensive results and recommendations.
364
+
365
+ ## Dependencies
366
+ Key dependencies (see [requirements.txt](requirements.txt) for complete list):
367
+ - `mlflow` - Experiment tracking and model registry
368
+ - `transformers` - HuggingFace transformer models
369
+ - `torch` or `tensorflow` - Deep learning backend
370
+ - `datasets` - HuggingFace datasets library
371
+ - `scikit-learn` - Classical ML algorithms
372
+ - `pandas`, `numpy` - Data manipulation
373
+ - `optuna` or `ray[tune]` - Hyperparameter optimization (optional)
374
+
375
+ ## License
376
+ [Add license information here]
377
+
378
+ ## Contact
379
+ [Add contact/team information here]