Text Classification
Transformers
Safetensors
Chinese
bert
ai-generated-text-detection
chinese
binary-classification
thesis
academic
Eval Results (legacy)
Instructions to use AnxForever/chinese-ai-detector-bert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AnxForever/chinese-ai-detector-bert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AnxForever/chinese-ai-detector-bert")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("AnxForever/chinese-ai-detector-bert") model = AutoModelForSequenceClassification.from_pretrained("AnxForever/chinese-ai-detector-bert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 8,432 Bytes
00e1a75 9423193 dfd7f4b 00e1a75 dfd7f4b 00e1a75 9423193 dfd7f4b 24e888e dfd7f4b 9423193 dfd7f4b 9423193 dfd7f4b f9727c4 dfd7f4b f9727c4 dfd7f4b f9727c4 dfd7f4b f9727c4 dfd7f4b f9727c4 dfd7f4b ea3fca0 dfd7f4b ea3fca0 dfd7f4b ea3fca0 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b ea3fca0 00e1a75 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 9423193 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 24e888e dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 24e888e 00e1a75 dfd7f4b 24e888e 00e1a75 dfd7f4b 00e1a75 24e888e dfd7f4b 00e1a75 ea3fca0 00e1a75 ea3fca0 00e1a75 ea3fca0 00e1a75 ea3fca0 dfd7f4b 00e1a75 ea3fca0 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b ea3fca0 dfd7f4b ea3fca0 9423193 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b 00e1a75 dfd7f4b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 | ---
language:
- zh
license: mit
library_name: transformers
pipeline_tag: text-classification
tags:
- ai-generated-text-detection
- chinese
- bert
- text-classification
- binary-classification
- thesis
- academic
base_model: bert-base-chinese
metrics:
- accuracy
- precision
- recall
- f1
model-index:
- name: chinese-ai-detector-bert-v11c
results:
- task:
type: text-classification
name: Text Classification
dataset:
type: AnxForever/chinese-ai-detection-dataset
name: Validation Set
split: validation
metrics:
- type: accuracy
value: 0.9875
name: Accuracy
- type: f1
value: 0.9883
name: F1
- task:
type: text-classification
name: Text Classification
dataset:
type: AnxForever/chinese-ai-detection-dataset
name: Independent Evaluation (910 samples)
split: test
metrics:
- type: accuracy
value: 0.9857
name: Accuracy
- type: f1
value: 0.9579
name: F1
- task:
type: text-classification
name: Text Classification
dataset:
type: AnxForever/chinese-ai-detection-dataset
name: Three-Set Average
metrics:
- type: accuracy
value: 0.9856
name: Accuracy
---
# Chinese AI-Generated Text Detector — BERT v11c (Boundary-Fix)
> 中文 AI 生成文本检测器(本科毕业设计最终版)
>
> A fine-tuned BERT model that classifies Chinese text as either **human-written (0)** or **AI-generated (1)**. The main released model is a document-level binary classifier; mixed-text boundary detection is an experimental extension provided by a separate span model.
---
## 📌 模型概述 / Overview
**中文**:本模型是基于 `bert-base-chinese` 微调的中文 AI 生成文本二分类器,为本科毕业设计「基于 BERT 微调的中文 AI 生成文本检测系统」的最终生产模型(v11c boundary-fix 版本)。当前主链路输出 **Human / AI 二分类**;`[SEP]` 边界标记与 Token 级 span detector 是配套的实验性扩展,用于探索构造型人机混写样本中的片段级分析。
**English**: A binary classifier fine-tuned on `bert-base-chinese` for Chinese AI-generated text detection. This is the final production checkpoint (v11c boundary-fix) of an undergraduate thesis project. `[SEP]` boundary markers and the token-level span detector are experimental extensions for constructed human/AI mixed-text analysis, not the default production inference path.
---
## 📊 评估指标 / Evaluation
| Dataset | Samples | Accuracy | Precision | Recall | F1 |
|---|---:|---:|---:|---:|---:|
| Validation set | 7,452 | **98.75 %** | 98.30 % | 99.37 % | 98.83 % |
| `core_v1_test_clean` | 545 | 97.98 % | 97.87 % | 98.77 % | 98.32 % |
| Independent eval (910) | 910 | **98.57 %** | 93.08 % | 98.67 % | 95.79 % |
| **Three-set average** | — | **98.56 %** | — | — | — |
> The metrics above evaluate the **document-level binary classifier**. The historical token-level boundary result belongs to the separate `chinese-ai-detector-span` experimental model and should not be mixed with the main classifier metrics.
### Independent eval by source (selected)
| Source | Samples | Accuracy |
|---|---:|---:|
| Toutiao News (all) | 377 | 100.0 % |
| Wikipedia CN | 119 | 99.16 % |
| formal_collected | 200 | 96.5 % |
| real_ai_gemini-3-pro-preview | 24 | 100.0 % |
| real_ai_deepseek-v3.2 | 8 | 100.0 % |
---
## 🏗️ 架构 / Architecture
- **Base model**: `bert-base-chinese` (12 layers, hidden 768, 12 heads, vocab 21,128)
- **Head**: `BertForSequenceClassification` (2 labels: `0 = human`, `1 = AI`)
- **Max sequence length**: 256 tokens (train), 512 (supported)
- **Framework**: `transformers 4.57.3`, PyTorch 2.0+
- **Parameters**: ~102M
### Training configuration
| Setting | Value |
|---|---|
| Base model | `bert-base-chinese` (via `bert_v7_improved` intermediate checkpoint) |
| Train samples | 63,113 |
| Validation samples | 7,452 |
| Epochs | 5 (best at epoch 2) |
| Batch size | 8 × 4 grad accum |
| Learning rate | 1e-5 |
| Label smoothing | 0.05 |
| Max length | 256 |
| Early stopping patience | 2 |
### Data changes vs. v10 baseline
- Removed 750 hard patterns + 1,767 unapproved samples + 7 length violations
- Added 300 formal-collected weak-domain samples
- Added 300 Llama-405B weak-domain samples
- Added **2,131 long-AI boundary-fix samples** (the key v11c contribution)
- Net change: +207 rows vs. v10
---
## 🚀 使用方法 / Usage
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
MODEL_ID = "AnxForever/chinese-ai-detector-bert"
TEMPERATURE = 0.8165 # Temperature scaling, calibrated on 910 samples (ECE=0.0034)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)
model.eval()
text = "这是一段需要检测的中文文本。"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
logits = model(**inputs).logits
# Apply temperature scaling for calibrated confidence
probs = torch.softmax(logits / TEMPERATURE, dim=-1)[0]
pred_idx = int(probs.argmax())
label = model.config.id2label[pred_idx] # "human-written" or "AI-generated"
print(f"{label} (confidence: {probs[pred_idx].item():.2%})")
```
### Label mapping
- `0` → human-written (人类撰写)
- `1` → AI-generated (AI 生成)
> **Note on Temperature Scaling**: `T = 0.8165` was calibrated on a held-out 910-sample set
> and brings ECE from 0.0121 down to **0.0034**. For uncalibrated probabilities, set `TEMPERATURE = 1.0`.
---
## 🎯 技术贡献 / Contributions
1. **Data-centric risk governance**
The v11c model keeps the BERT backbone fixed and improves robustness through data cleaning, weak-domain supplementation, long-AI supplementation, and calibrated inference.
2. **`[SEP]` boundary-marker experiment**
In constructed C2-style mixed samples, `[SEP]` was used as an explicit boundary hint between known human and AI segments. This is an engineering experiment for mixed-text modeling, not a claim that `[SEP]` itself can identify authorship without labels.
3. **Two-stage experimental extension**
- Stage 1: this model — document-level Human / AI classification
- Stage 2: separate span detector — token-level Human / AI tagging on mixed-text samples
- See [`AnxForever/chinese-ai-detector-span`](https://huggingface.co/AnxForever/chinese-ai-detector-span)
4. **Long-AI boundary-fix (v11c)**
针对长 AI 段落在边界处易被误判的问题,补充 2,131 条长 AI 边界样本,使 256+ token 桶的准确率恢复到 V10 水平。
### Note on mixed-text boundary detection
The boundary module was trained on a relatively small constructed mixed-text set. It is useful for demonstration, teaching, and secondary development, but it should be treated as an experimental prototype. For real business scenarios, mixed human/AI data from the target domain should be collected, labeled, retrained, and evaluated before deployment.
---
## ⚠️ 局限性 / Limitations
- 仅针对**中文**文本;对英文或其他语言无保证。
- 训练语料偏新闻/百科/技术/正式文体,对**诗歌、古文、社交媒体短文本**可能欠拟合。
- 当前默认发布能力是**篇章级二分类**;人机混写边界定位属于实验性扩展,不建议直接作为商业审核结论。
- 训练数据主要来自 DeepSeek、Gemini、GPT、Llama-405B 等主流模型;对**经过重度改写**的 AI 文本仍有遗漏风险。
- 对短文本、强人工改写文本、多次交替混写文本和目标域外文本,不保证固定准确率。
---
## 🗂️ 相关资源 / Related
- 📊 训练数据集 / Dataset: [`AnxForever/chinese-ai-detection-dataset`](https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset)
- 🎯 边界检测器 / Span detector: [`AnxForever/chinese-ai-detector-span`](https://huggingface.co/AnxForever/chinese-ai-detector-span)
---
## 📜 License
MIT License
## ✍️ Citation
```bibtex
@misc{anxforever2026chineseaidetectorbert,
title = {Chinese AI-Generated Text Detector with Boundary Markers (BERT v11c)},
author = {AnxForever},
year = {2026},
howpublished = {\url{https://huggingface.co/AnxForever/chinese-ai-detector-bert}},
note = {Undergraduate thesis project}
}
```
|