Instructions to use sivasub987/iso20022-extract-53m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sivasub987/iso20022-extract-53m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sivasub987/iso20022-extract-53m", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("sivasub987/iso20022-extract-53m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sivasub987/iso20022-extract-53m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sivasub987/iso20022-extract-53m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/sivasub987/iso20022-extract-53m
- SGLang
How to use sivasub987/iso20022-extract-53m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sivasub987/iso20022-extract-53m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sivasub987/iso20022-extract-53m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use sivasub987/iso20022-extract-53m with Docker Model Runner:
docker model run hf.co/sivasub987/iso20022-extract-53m
Resolving role consistency and memorization bottlenecks via 2-pass block recycling
Hi Siva,
Documenting the failure modes of deceptive validation metrics (the uninitialized RoPE inv_freq buffer and teacher-forcing loops) is one of the most transparent and insightful write-ups on Hugging Face.
Looking at your architecture trade-offs and the C8 (role consistency = 0.0) / C5 (field accuracy = 71.7%) gates:
- Parameter capacity vs training token exposure:
- 3,800 documents across 256 tokens for 3 epochs equals ~2.9M tokens of total exposure.
- At 53M parameters, the network has ~18x more weights than training tokens. As you noted regarding the 97.5% vs 61.2% split, 53M parameters easily memorize synthetic fixtures rather than learning generalized extraction transforms.
- Role inversion in unlabelled running prose:
Disentangling debtor from creditor without canonical field tags requires multi-step syntactic binding: first identifying entity spans (IBANs, names, amounts), then traversing prepositional phrases ("on behalf of", "favoring") to assign semantic roles. With d_model=512 and intermediate FFN=768, 20 shallow layers struggle with this multi-hop resolution.
In an open architecture project called Maba (101M reference model: https://huggingface.co/AndrewThompson1233/maba-v1-architecture), we handle this regime using deterministic 2-pass block recycling:
Reduce physical depth from 20 blocks to 10, and cycle representations through them twice during the forward pass.
Condition each pass using Split RMSNorm (distinct scale vectors for pass 0 and pass 1). Pass 0 handles token-level span detection, while pass 1 refines relational role assignment.
This cuts physical weights from 53M down to ~33M parameters, curbing the capacity for verbatim memorization on small corpora while preserving 20 effective layers of relational depth.
Compute and memory fit effortlessly within the 4 GB VRAM ceiling of mobile GPUs like your RTX 2050.
Did you test whether scaling down physical layer depth or varying parameter-to-token ratios improved generalization on unseen hostile-difficulty draws?
Best,
Andrew