Instructions to use ibm-granite/granitelib-rag-gpt-oss-r1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Granite Library
How to use ibm-granite/granitelib-rag-gpt-oss-r1.0 with Granite Library:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Download hallucination_detection/README.md from ibm-granite/granitelib-rag-gpt-oss-r1.0: direct link, hf CLI and curl.
- Browser
- Download file 7.82 kB
-
https://huggingface.co/ibm-granite/granitelib-rag-gpt-oss-r1.0/resolve/664dfb1f82f4d35b80c3eaed72759b1ba5befbe5/hallucination_detection/README.md
- Command line
-
hf download hf://ibm-granite/granitelib-rag-gpt-oss-r1.0@664dfb1f82f4d35b80c3eaed72759b1ba5befbe5/hallucination_detection/README.md
-
curl -L -o README.md https://huggingface.co/ibm-granite/granitelib-rag-gpt-oss-r1.0/resolve/664dfb1f82f4d35b80c3eaed72759b1ba5befbe5/hallucination_detection/README.md
Intrinsics for Hallucination Detection
Model Summary
This is a RAG-specific family of intrinsics fine-tuned for the hallucination detection task. Given a multi-turn conversation between a user and an AI assistant, ending with an assistant response and a set of documents/passages on which the last assistant response is supposed to be based, the adapter outputs a hallucination label (faithful/partial/unfaithful/NA) for each sentence in the assistant response.
We provide two intrinsics implemented as LoRA adapters trained over Granite-3.3-2b-instruct and Granite-3.3-8b-instruct, respectively.
- Developer: IBM Research
- Model type: LoRA adapter for ibm-granite/granite-3.3-2b-instruct and ibm-granite/granite-3.3-8b-instruct
- License: Apache 2.0
Intended use
This is a family of hallucination detection intrinsics that gives the ability to identify hallucination risks for the sentences in the last assistant response in a multi-turn RAG conversation based on a set of provided documents/passages.
Note: While you can invoke the hallucination detection intrinsic directly, it is strongly recommended to call it through granite-common, which wraps the model with a tailored I/O processor, enabling a friendlier development interface. The I/O processor takes care of several data transformation/validation tasks that would be otherwise required (incl. splitting the input documents and assistant response into sentences before calling the intrinsic as well as validating the intrinsic's output). We next describe the input/output of the hallucination detection intrinsics when invoked through granite-common.
Intrinsic input: The hallucination detection intrinsic takes as input an OpenAI-compatible chat completion request. This request includes: a list of conversation turns that ends with the assistant’s response (the response to be checked for hallucinations) and a list of reference documents that the final assistant response should be grounded on. See the code snippets in the Quickstart Example section below for examples of how to format the chat completion request as a JSON object.
Intrinsic output: The output of the hallucination detection intrinsic is formatted as the result of the original chat completion request containing the hallucinations detected for the last assistant response. The hallucinations are provided in the form of a JSON array, whose items include the text and begin/end of a response span (sentence) together with the text, faithfulness_likelihood of the response sentence, and the explanation for the faithfulness_likelihood.
Going from input to output: When calling the intrinsic through granite-common one should follow the steps below to transform the intrinsic input to the corresponding output. These steps are also exemplified in the code snippets included in the Quickstart Example section below. Given an input chat completion request, the request should be passed to the corresponding input processor (also referred to as IntrinsicsRewriter) provided by granite-common. The input processor converts the request to the appropriate format expected by the underlying hallucination detection model. This includes, among others, splitting the last assistant response and the documents into sentences and prepending them with sentence IDs as well as introducing an appropriate task-specific instruction. The input processor's result should then be passed to the underlying hallucination detection model for inference. The model identifies hallucinations using a compact representation consisting of sentence IDs in the last assistant response and documents. This output should finally be passed to the appropriate output processor (also referred to as IntrinsicsResultProcessor) provided by granite-common. The output processor converts the low-level raw model output to the final output by, among others, mapping the sentence IDs back to response and document spans. The result is an application-friendly format ready for consumption by downstream applications.
Quickstart Example
The recommended way to call this intrinsic is through the Mellea framework. Here is some example code for calling this intrinsic from Mellea:
from mellea.backends.huggingface import LocalHFBackend
from mellea.stdlib.base import ChatContext, Document
from mellea.stdlib.chat import Message
from mellea.stdlib.intrinsics import rag
import json
backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct")
context = (
ChatContext()
.add(Message("assistant", "Hello there, how can I help you?"))
.add(Message("user", "Tell me about some yellow fish."))
)
assistant_response = "Purple bumble fish are yellow. Green bumble fish are also yellow."
documents = [
Document(
doc_id="1",
text="The only type of fish that is yellow is the purple bumble fish.",
)
]
result = rag.flag_hallucinated_content(assistant_response, documents, context, backend)
print(f"Result of hallucination check: {json.dumps(result, indent=2)}")
Training Details
The process of generating the training data for the hallucination detection intrinsic consisted of two main steps:
Multi-turn RAG conversation generation: Starting from publicly available document corpora, we generated a set of multi-turn RAG data, consisting of multi-turn conversations grounded on passages retrieved from the corpus. For details on the RAG conversation generation process, please refer to the Granite Technical Report and Lee, Young-Suk, et al..
Faithfulness label generation: For creating the faithfulness labels for the responses, we used a multi-step synthetic hallucination label and reasoning generation pipeline. This process resulted in ~50K data instances, which were used to train the LoRA adapter.
Training Data
The following public datasets were used as seed datasets for the multi-turn RAG conversation generation process:
CoQA - Wikipedia passages
Evaluation
We evaluated the LoRA adapter on the QA portion of the RAGTruth benchmark. We compare the response-level hallucination detection performance between the LoRA adapter and the methods reported in the RAGTruth paper. The responses that obtain a faithfulness labels partial or unfaithful for at least one sentence are considered as hallucinated responses.
The results are shown in the table below. The results for the baselines are extracted from the RAGTruth paper.
| Model | Precision | Recall | F1 |
|---|---|---|---|
| GPT 4o mini (prompted) | 46.8 | 59.6 | 52.4 |
| GPT 4o (prompted) | 49.5 | 60.1 | 54.3 |
| gpt-4-turbo (prompted) | 33.2 | 90.6 | 45.6 |
| SelfCheckGPT | 35.0 | 58.0 | 43.7 |
| LMvLM | 18.7 | 76.9 | 30.1 |
| Granite 3.3-2b_hallucination-detection_LoRA | 55.8 | 74.9 | 63.9 |
| Granite 3.3-8b_hallucination-detection_LoRA | 58.1 | 77.6 | 66.5 |