thefinalboss commited on
Commit
454e3ba
·
verified ·
1 Parent(s): 1f1d55d

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +5 -4
README.md CHANGED
@@ -30,7 +30,7 @@ datasets:
30
  ![PyTorch](https://img.shields.io/badge/PyTorch-2.9-orange)
31
  ![Params](https://img.shields.io/badge/params-1.05B-red)
32
  ![Status](https://img.shields.io/badge/status-active-brightgreen)
33
- ![Datasets](https://img.shields.io/badge/datasets-3.25B%20tokens-purple)
34
 
35
  ---
36
 
@@ -73,20 +73,21 @@ Fractus is a **Continuous Cognitive Agent** — an AI that works like a brain, n
73
 
74
  ---
75
 
76
- ## Datasets (3.25 Billion Tokens)
77
 
78
  Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets):
79
 
80
  | Dataset | Tokens | Content |
81
  |---|---|---|
82
  | **neuro-paradigms-1b** | **1.78B** | 300 neuroscience → software architecture paradigms |
 
83
  | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
84
  | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
85
  | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
86
- | **all-github-repos** | **54M** | 49 of your GitHub repos (self-aware codebase) |
87
  | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
 
88
  | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
89
- | **Total** | **~3.25B** | |
90
 
91
  The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code.
92
 
 
30
  ![PyTorch](https://img.shields.io/badge/PyTorch-2.9-orange)
31
  ![Params](https://img.shields.io/badge/params-1.05B-red)
32
  ![Status](https://img.shields.io/badge/status-active-brightgreen)
33
+ ![Datasets](https://img.shields.io/badge/datasets-4.15B%20tokens-purple)
34
 
35
  ---
36
 
 
73
 
74
  ---
75
 
76
+ ## Datasets (4.15 Billion Tokens)
77
 
78
  Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets):
79
 
80
  | Dataset | Tokens | Content |
81
  |---|---|---|
82
  | **neuro-paradigms-1b** | **1.78B** | 300 neuroscience → software architecture paradigms |
83
+ | **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms |
84
  | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
85
  | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
86
  | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
 
87
  | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
88
+ | **all-github-repos** | **54M** | 49 of your GitHub repos (self-aware codebase) |
89
  | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
90
+ | **Total** | **~4.15B** | |
91
 
92
  The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code.
93