Acacia-315M

Paper: https://arxiv.org/abs/2609.30894

Just provide a graph and a few examples. Acacia is a graph foundation model for node classification, clustering, link prediction, graph generation, and more without additional training.

Suppose you have citation links between papers and features for each paper. Label some of the papers with their research fields and provide them to Acacia, and it can predict labels for the remaining papers. It can also group similar nodes without being given any labels at all.

Acacia was trained on the Common Crawl web graph, where web pages are nodes and hyperlinks are edges. A single model with approximately 315 million parameters can be used on graphs with different feature dimensionalities and label types. There is no need to train an additional classifier or feature transformation module for each new graph.

  • Use it without updating weights: Acacia supports In-Context Learning (ICL), making predictions from inputs that include labeled examples.
  • Use it without labels: Nodes can be assigned to groups sequentially without specifying the number of clusters in advance.
  • Use different features: Your numerical features are converted to 64 dimensions using a random projection that requires no training, then provided as input.
  • Trained from scratch on graphs: Acacia was trained on approximately 11 billion graph tokens in total, without using pretrained language model weights.

The inputs are numerical node features, edges, and, when needed, known labels.

Item Description
Model name Acacia-315M
Parameters 315,023,040 (approximately 315M)
Architecture Decoder Transformer, 32 layers, hidden dimension 960
Pretraining data Graphs constructed from Common Crawl web pages, links, titles, and body text
Training volume Approximately 10 billion tokens with title features + approximately 1 billion tokens incorporating body features and ICL
Implementation Custom PyTorch model GraphLM, inference weights model.pt
Paper Training Graph Foundation Models with The Web Graph

What can it do?

The same model can be used for different tasks by changing the information included in the input and what it is asked to generate.

Task Information provided as input Output
Node classification A graph and labels for some of its nodes A predicted label for a specified node
Classification with additional examples (ICL) The target graph and labeled examples placed in separate connected components Predictions based on the example labels
Unsupervised clustering The graph alone; neither ground-truth labels nor the number of clusters is required A cluster ID for each node
Link prediction and graph generation The parts to condition on, such as nodes and edges A continuation consisting of edges, nodes, and so on

Training and how the model works

Acacia is an autoregressive model that converts a graph into node, edge, and node-label records and predicts their continuation. Each record is preceded by a header specifying the target node and the type to generate. This allows the same model to be instructed to, for example, predict the label of a particular node or generate an edge connected to that node.

Node IDs and label IDs use random codes assigned for each input. The model is designed to learn relationships within that input rather than memorize specific class numbers or feature column indices. When a new group is needed, it selects NEW and assigns a new label code.

During training, small graphs are extracted by following links between web pages, and page domains serve as labels. Titles and body text are converted into 64-dimensional hashed features.

Evaluation results

The following results are those reported in the paper. All evaluations keep Acacia's weights frozen, with no additional training on downstream tasks. "Untrained" refers to a randomly initialized model with the same architecture, whose weights are also frozen. Mean ± sample standard deviation is reported across 3 seeds.

The datasets are the Cora, CiteSeer, and PubMed citation graphs and the ogbn-products product graph. For classification, we use the Planetoid public splits for the former and the official OGB split for the latter.

Node classification: predicting from known labels in the neighborhood

For each target node, we extract a neighborhood of up to 128 nodes and provide the training labels naturally present in it. No separate example graphs are added. Results are accuracy (%).

Model / baseline Cora CiteSeer PubMed ogbn-products
Chance level 41.94 ± 0.29 32.38 ± 0.13 19.68 ± 0.72 44.15 ± 0.66
Untrained, frozen 41.07 ± 1.37 33.07 ± 0.67 19.80 ± 0.56 45.00 ± 0.17
Acacia, frozen 58.17 ± 1.35 39.50 ± 1.22 19.87 ± 0.74 66.00 ± 0.72

We evaluate 1,000 test nodes per dataset. For ogbn-products, these are a subset sampled from the official test set. Candidate labels are restricted to those appearing in the input, so cases where the correct class label is absent from the neighborhood or the candidate set is empty are also counted as errors. The chance level is the expected accuracy of selecting uniformly from these candidates.

Acacia achieves higher accuracy than the untrained model on Cora, CiteSeer, and ogbn-products. PubMed, however, has few labeled nodes, and its results are close to chance in this setting.

ICL: adding labeled examples as separate components

We add 1-2 training examples per class as separate connected components to the target neighborhood, which contains up to 32 nodes. The entire input, including examples, has at most 128 nodes. Results are accuracy (%).

Model / baseline Cora CiteSeer PubMed ogbn-products
Chance level 14.29 16.67 33.33 2.35
Untrained, frozen 15.27 ± 0.92 17.30 ± 1.57 33.67 ± 1.03 2.37 ± 0.40
Acacia, frozen 54.33 ± 1.61 39.37 ± 2.06 37.43 ± 2.05 68.38 ± 0.32

We evaluate 1,000 test nodes each for Cora, CiteSeer, and PubMed, and 2,000 nodes sampled from the official test set for ogbn-products. Predictions are selected from the labels presented in the input. The chance level for ogbn-products accounts for the candidate classes with training examples and the coverage of the true classes.

Acacia outperforms both the untrained model and the chance level on all 4 datasets. This shows that a model trained on the web graph can handle other graphs and their examples without changing its weights.

Unsupervised clustering: no ground-truth labels or cluster count provided

We provide only the graph and sequentially generate cluster assignments for every input node. "1 root" collects up to 128 nodes by BFS from 1 node. "4 roots" collects up to 32 nodes from each of 4 nodes and uses their union. A graph constructed from 4 roots does not necessarily have 4 separate components.

The metrics are ARI (Adjusted Rand Index) and AMI (Adjusted Mutual Information). For both, larger values indicate greater agreement with the ground-truth grouping, and 1 indicates a perfect match. For the evaluation cases below, the baseline for agreement by chance is 0.

Setting / metric Model Cora CiteSeer PubMed ogbn-products
1 root / ARI Untrained -0.0013 ± 0.0029 0.0050 ± 0.0167 0.0006 ± 0.0047 0.0009 ± 0.0023
1 root / ARI Acacia 0.1191 ± 0.0084 0.0667 ± 0.0304 0.0387 ± 0.0068 0.0047 ± 0.0116
1 root / AMI Untrained -0.0009 ± 0.0046 0.0096 ± 0.0116 0.0039 ± 0.0011 0.0001 ± 0.0039
1 root / AMI Acacia 0.1437 ± 0.0150 0.0792 ± 0.0341 0.0599 ± 0.0035 0.0061 ± 0.0099
4 roots / ARI Untrained 0.0013 ± 0.0039 0.0008 ± 0.0036 0.0009 ± 0.0030 0.0016 ± 0.0015
4 roots / ARI Acacia 0.2643 ± 0.0120 0.2833 ± 0.0311 0.2034 ± 0.0043 0.0117 ± 0.0069
4 roots / AMI Untrained 0.0006 ± 0.0024 0.0025 ± 0.0046 0.0003 ± 0.0018 0.0006 ± 0.0006
4 roots / AMI Acacia 0.3295 ± 0.0053 0.3241 ± 0.0178 0.2627 ± 0.0121 0.0222 ± 0.0106

Results are aggregated over subgraphs with at least 2 ground-truth classes, excluding cases where every node belongs to a different class. Ground-truth labels are used only to determine eligibility and compute scores, never as inputs or during generation. We first average the scores across graphs within each seed, then report the mean and sample standard deviation of those 3 averages.

The model exhibits an ability to group nodes in citation graphs even without labels.

Citation

@article{sato2026acacia,
  author = {Ryoma Sato},
  title  = {Training Graph Foundation Models on The Web Graph},
  year   = {2026}
  doi    = {10.48550/arXiv.2609.30894},
  url    = {https://arxiv.org/abs/2609.30894},
}
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for joisino/acacia-315m