anlp-neutral-led

LED-base checkpoints for neutral article generation (IIIT Hyderabad ANLP project: Event-Aware Representation Learning for Neutral Multi-Source News). The model gets one partisan news article and writes the same article without framing. Each checkpoint lives in <experiment>/<run>/.

Runs

subfolder negatives triplet 位 best epoch test F1 test gen loss test token acc length ratio
side-neutral-full-publisher/lambda01 publisher negatives 0.1 3 0.811 0.434 0.890 1.02
side-neutral-full-publisher/lambda02 publisher negatives 0.2 3 0.810 0.430 0.890 1.01
side-neutral-full-publisher/lambda035 publisher negatives 0.35 3 0.809 0.435 0.889 1.02
side-neutral-full/lambda0 ideology negatives 0 3 0.810 0.430 0.890 1.02
side-neutral-full/lambda01 ideology negatives 0.1 3 0.810 0.434 0.890 1.02
side-neutral-full/lambda02 ideology negatives 0.2 3 0.811 0.435 0.890 1.02
side-neutral-full/lambda035 ideology negatives 0.35 3 0.810 0.435 0.889 1.02

Test set: 1,775 held-out articles. F1 is whitespace-token overlap with the neutral reference after greedy decoding (max 512 new tokens); token accuracy is teacher-forced. 位 = 0 never uses negatives, so it is the control for every experiment.

Experiments

All runs share the data and training setup:

  • Data: left/right_neutral_dedup (BigNewsAlign articles, neutral targets generated by Nemotron-3-Ultra), restricted to articles that share an event with an opposite-side article: 12,484 train / 1,778 val / 1,775 test, no event, article or body shared across splits (dataset/side_neutral_full on the data branch).
  • Objective: generation loss + 位 脳 triplet margin loss (margin 0.2) on L2-normalised masked-mean encoder states. Anchor = input article, positive = opposite-side article about the same event; the negative depends on the experiment.
  • Training: fp32, lr 3e-5, effective batch 8, 4 epochs (best by validation loss), 10% warmup, max 512 input / 512 target tokens, attention window 512, seed 42, one T4 per run.

side-neutral-full

Negative = same-side article from a different event group (proposal variant b; it shares the anchor's publisher ~30% of the time, since each publisher is on one side).

side-neutral-full-publisher

Negative = article from the anchor's own publisher, different event group (proposal variant a). Anchors, positives and targets are identical to side-neutral-full.

Loading

from transformers import AutoTokenizer, LEDForConditionalGeneration

sub = "side-neutral-full/lambda035"
tok = AutoTokenizer.from_pretrained("RaunakSeksaria/anlp-neutral-led", subfolder=sub)
model = LEDForConditionalGeneration.from_pretrained("RaunakSeksaria/anlp-neutral-led", subfolder=sub)

source = f"{title} {tok.sep_token} {body}"   # training input format
enc = tok(source, return_tensors="pt", truncation=True, max_length=512)
glob = enc["attention_mask"].new_zeros(enc["attention_mask"].shape)
glob[:, 0] = 1                                   # global attention on the first token
out = model.generate(**enc, global_attention_mask=glob, num_beams=1,
                     do_sample=False, max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))

Article embeddings (what the triplet loss shapes): mean of model.get_encoder()(...).last_hidden_state over non-padding tokens, then L2 normalisation. metrics/ in each subfolder holds the training history, config and validation/test metrics; the repo is private.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for RaunakSeksaria/anlp-neutral-led

Finetuned
(45)
this model