diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000000000000000000000000000000000000..cafbd372087e6a6a45eb56dd5a739e1fe8bcbc29 --- /dev/null +++ b/.gitignore @@ -0,0 +1,18 @@ +# Dependencies +node_modules/ +venv/ + +# Python +__pycache__/ +*.pyc +*.egg-info/ + +# Environment +.env + +# Data / generated files +saved_articles.json +benchmark_results/ + +# Next.js build output +.next/ diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000000000000000000000000000000000000..ec7114f5fc5515947e8c9b352216c6c6665a9e2f --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,106 @@ +# CLAUDE.md + +This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. + +## Project Overview + +TruthScan AI is a master's thesis full-stack application for real-time news credibility analysis. It fetches articles from RSS feeds, runs them through HuggingFace Transformer models (sentiment analysis + fake news detection), and presents results in a bilingual (PL/EN) Next.js dashboard. + +## Commands + +### Backend (`TruthScan AI_backend/`) + +```bash +# Install dependencies +cd "TruthScan AI_backend" +pip install -r requirements.txt + +# Start dev server (http://127.0.0.1:8000) +python -m uvicorn app.main:app + +# Run tests +python truthscan_test.py +``` + +Swagger UI available at `http://localhost:8000/docs`. + +### Frontend (`TruthScan AI_frontend/`) + +```bash +# Install dependencies +cd "TruthScan AI_frontend" +npm install + +# Start dev server (http://localhost:3000) +npm run dev + +# Build for production +npm run build + +# Lint +npm run lint +``` + +## Architecture + +### Data Flow + +``` +Browser (Next.js :3000) + → HTTP/SSE → FastAPI (:8000) + → RSS fetch → feedparser + BeautifulSoup + → NLP pipeline → HuggingFace Transformers + → JSON/SSE response → React UI +``` + +### Backend (`app/`) + +- **`config.py`** — All constants: 10 RSS sources (BBC, CNN, NYTimes, Guardian, AlJazeera, PolsatNews, etc.), CORS settings, sentiment label mappings (PL/EN), cache TTLs. +- **`nlp_service.py`** — Plug-in NLP pipeline: + - Abstract base class `ModelAdapter` with `analyze_sentiment()` and `analyze_fake_news()`. + - Three adapters: `RoBERTaAdapter` (en), `XLMRoBERTaAdapter` (pl/en/no), `NorBERTAdapter` (no). + - All adapters lazy-load their pipelines on first use. + - Global registry `_REGISTRY`; active adapter changed via `set_active_adapter(name)`. + - `register_adapter(adapter)` adds custom/fine-tuned checkpoints at runtime. + - Public functions `analyze_news()` / `analyze_news_batch()` preserve the original interface — `routes/news.py` requires no changes. + - Uses `ThreadPoolExecutor` (max 3 workers) for parallel batch processing. +- **`rss_utils.py`** — RSS fetching with 7s timeout, HTML stripping via BeautifulSoup, simple in-memory TTL cache. +- **`storage.py`** — Thread-safe JSON file persistence for saved articles (`saved_articles.json`). +- **`routes/news.py`** — Main endpoints: `GET /news/{source}` (5 articles with NLP), `GET /stream-news/{source}` (SSE streaming), `GET /emotion-stats/{source}`, `GET /charts-data`. +- **`routes/saved.py`** — CRUD for saved articles. +- **`routes/misc.py`** — `GET /sources` lists available RSS sources. + +### Frontend (`app/` + `components/` + `lib/`) + +- **Routing**: Next.js App Router — pages at `app/page.tsx` (home), `app/dashboard/page.tsx`, `app/saved/page.tsx`. +- **State**: Zustand store in `stores/newsCache.tsx` — caches articles per source/language with 5-min TTL, persisted to `localStorage` (key: `truthscan_news_cache_v1`). +- **API client**: `lib/fetchNews.ts` — `fetchOneSource()` and `fetchAllNews()` (concurrent, 3 workers default), plus `normalizeArticle()`. +- **SSE streaming**: `hooks/useNewsStream.ts` consumes `GET /stream-news/{source}`, tracks progress and collects articles, writes to Zustand cache on completion. +- **i18n**: `lib/locales.js` holds all PL/EN UI strings. Language state managed in `hooks/useLanguage.ts`, synced via `localStorage` and `app:langchange` custom event. +- **Charts**: Recharts via `hooks/useDashboardCharts.ts` and `hooks/useCachedEmotionStats.ts`. +- **PDF export**: `hooks/usePDFExport.ts` (jsPDF) + `components/PDFGenerator.tsx` (react-to-print). +- **Dark mode**: Tailwind `.dark` class toggle, CSS variables in `styles/globals.css`. + +## Master's Thesis Goals + +The thesis extends TruthScan AI by comparing NLP models across three languages: **Polish**, **Norwegian**, and **English**. Models under comparison: **XLM-RoBERTa**, **NorBERT 3**, **HerBERT**. + +Planned work items: + +1. **`nlp_service.py` plug-in architecture** — replace the hardcoded model with a swappable interface so each model (XLM-RoBERTa, NorBERT 3, HerBERT) can be loaded and hot-swapped without changing route logic. + +2. **Norwegian RSS sources in `config.py`** — add NRK, VG, and Dagbladet alongside the existing 10 sources. + +3. **`SENTIMENT_MAP` extension** — add a `"no"` (Norwegian) key to the sentiment label mapping in `config.py`, parallel to the existing `"pl"` and `"en"` keys. + +4. **Benchmarking module** — new module that measures **F1**, **accuracy**, and **inference time** per model/language combination and exposes results via a dedicated API endpoint (or offline report). + +5. **PostgreSQL migration** — replace `storage.py` / `saved_articles.json` with a PostgreSQL backend (SQLAlchemy or asyncpg). `storage.py` read/write interface should be preserved so routes need minimal changes. + +6. **Public deployment** — frontend on **Vercel**, backend + models on **Hugging Face Spaces**. + +### Environment + +- Backend API base URL: `process.env.NEXT_PUBLIC_API_URL` (defaults to `http://127.0.0.1:8000`). +- Backend CORS is open (`"*"`) — intentional for thesis/dev use. +- NLP models are downloaded automatically by HuggingFace on first run (can be slow). diff --git a/TruthScan AI_backend/app/__init__.py b/TruthScan AI_backend/app/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..e69de29bb2d1d6434b8b29ae775ad8c2e48c5391 diff --git a/TruthScan AI_backend/app/benchmark.py b/TruthScan AI_backend/app/benchmark.py new file mode 100644 index 0000000000000000000000000000000000000000..203ce3fbcb658b3a4cbb4837152748873373e8b1 --- /dev/null +++ b/TruthScan AI_backend/app/benchmark.py @@ -0,0 +1,240 @@ +""" +Moduł benchmarkowania modeli NLP. + +Mierzy czas inferencji, rozkład sentymentów i prawdopodobieństwo fake news +dla każdego adaptera (roberta, xlm-roberta, norbert) na próbce tekstów +z trzech grup językowych (en, pl, no). + +Użycie standalone: + python -m app.benchmark + +Użycie z API: + GET /benchmark + GET /benchmark?adapters=roberta,xlm-roberta&langs=en,pl +""" + +import csv +import json +import time +from collections import Counter +from pathlib import Path +from typing import Dict, List, Optional + +from .nlp_service import get_adapter, ModelAdapter, analyze_news + +# --------------------------------------------------------------------------- +# Próbka tekstów testowych +# --------------------------------------------------------------------------- + +SAMPLE_TEXTS: Dict[str, List[str]] = { + "en": [ + "The government announced new economic reforms to boost growth and reduce unemployment.", + "Flooding devastated coastal towns overnight, leaving thousands homeless.", + "Scientists discover a new vaccine that shows 95% efficacy against the virus.", + "Stock markets surged to record highs after positive inflation data.", + "A major scandal erupted as leaked documents exposed corporate corruption.", + "The peace talks collapsed after both sides failed to reach an agreement.", + "Renewable energy investments hit an all-time high this quarter.", + "Crime rates in the capital have dropped significantly over the past year.", + ], + "pl": [ + "Rząd ogłosił nowe reformy gospodarcze mające na celu pobudzenie wzrostu.", + "Powódź zniszczyła nadmorskie miejscowości, tysiące osób zostało bez dachu.", + "Naukowcy odkryli szczepionkę o 95-procentowej skuteczności przeciw wirusowi.", + "Giełda osiągnęła rekordowe poziomy po pozytywnych danych o inflacji.", + "Wybuchł wielki skandal po ujawnieniu dokumentów o korupcji korporacyjnej.", + "Rozmowy pokojowe załamały się po niepowodzeniu negocjacji.", + "Inwestycje w energię odnawialną osiągnęły historyczny rekord w tym kwartale.", + "Wskaźniki przestępczości w stolicy znacząco spadły w ciągu ostatniego roku.", + ], + "no": [ + "Regjeringen kunngjorde nye økonomiske reformer for å øke veksten.", + "Flom ødela kystbyer over natten og etterlot tusenvis uten hjem.", + "Forskere oppdaget en vaksine med 95 prosent effektivitet mot viruset.", + "Aksjemarkedene steg til rekordhøyder etter positive inflasjonsdata.", + "En stor skandale brøt ut da lekkede dokumenter avslørte korrupsjon.", + "Fredssamtalene brøt sammen etter at begge sider ikke klarte å bli enige.", + "Investeringer i fornybar energi nådde en historisk topp dette kvartalet.", + "Kriminalitetsratene i hovedstaden har falt betydelig det siste året.", + ], +} + +# --------------------------------------------------------------------------- +# Typy wyników +# --------------------------------------------------------------------------- + +BenchmarkResult = Dict # TypedDict zastąpiony zwykłym Dict dla czytelności + +# --------------------------------------------------------------------------- +# Funkcje benchmarkowania +# --------------------------------------------------------------------------- + +def _run_single( + adapter: ModelAdapter, + text: str, + lang: str, +) -> Dict: + """Uruchamia analyze_news dla jednego tekstu i mierzy czas.""" + start = time.perf_counter() + result = analyze_news(text, lang=lang, adapter=adapter) + elapsed_ms = (time.perf_counter() - start) * 1000 + return {**result, "inference_time_ms": elapsed_ms} + + +def run_benchmark( + adapter_names: Optional[List[str]] = None, + langs: Optional[List[str]] = None, +) -> List[BenchmarkResult]: + """ + Uruchamia benchmark dla wskazanych adapterów i języków. + + Args: + adapter_names: Lista nazw adapterów; None = wszystkie trzy. + langs: Lista kodów języków; None = ['en', 'pl', 'no']. + + Returns: + Lista słowników z wynikami — jeden wpis na kombinację adapter × język. + """ + if adapter_names is None: + adapter_names = ["roberta", "xlm-roberta", "norbert"] + if langs is None: + langs = ["en", "pl", "no"] + + results: List[BenchmarkResult] = [] + + for adapter_name in adapter_names: + try: + adapter = get_adapter(adapter_name) + except ValueError as exc: + results.append({ + "adapter_name": adapter_name, + "error": str(exc), + }) + continue + + for lang in langs: + texts = SAMPLE_TEXTS.get(lang, []) + if not texts: + continue + + per_text: List[Dict] = [] + for text in texts: + try: + per_text.append(_run_single(adapter, text, lang)) + except Exception as exc: + per_text.append({ + "sentiment": None, + "fake_probability": None, + "sentiment_score": None, + "inference_time_ms": None, + "error": str(exc), + }) + + # Agregacja + valid = [r for r in per_text if r.get("inference_time_ms") is not None] + times = [r["inference_time_ms"] for r in valid] + fakes = [r["fake_probability"] for r in valid if r.get("fake_probability") is not None] + sentiments = [r["sentiment"] for r in valid if r.get("sentiment")] + + results.append({ + "adapter_name": adapter_name, + "language": lang, + "sample_size": len(texts), + "successful_runs": len(valid), + "avg_inference_time_ms": round(sum(times) / len(times), 2) if times else None, + "min_inference_time_ms": round(min(times), 2) if times else None, + "max_inference_time_ms": round(max(times), 2) if times else None, + "avg_fake_probability": round(sum(fakes) / len(fakes), 2) if fakes else None, + "sentiments_distribution": dict(Counter(sentiments)), + "per_text": per_text, + }) + + return results + + +# --------------------------------------------------------------------------- +# Eksport wyników +# --------------------------------------------------------------------------- + +def export_json(results: List[BenchmarkResult], path: Path) -> None: + """Zapisuje pełne wyniki (z per_text) do pliku JSON.""" + path.parent.mkdir(parents=True, exist_ok=True) + with open(path, "w", encoding="utf-8") as fh: + json.dump(results, fh, ensure_ascii=False, indent=2) + + +def export_csv(results: List[BenchmarkResult], path: Path) -> None: + """ + Zapisuje wyniki zbiorcze (bez per_text) do pliku CSV. + Jeden wiersz = jedna kombinacja adapter × język. + """ + path.parent.mkdir(parents=True, exist_ok=True) + summary_fields = [ + "adapter_name", "language", "sample_size", "successful_runs", + "avg_inference_time_ms", "min_inference_time_ms", "max_inference_time_ms", + "avg_fake_probability", "sentiments_distribution", + ] + with open(path, "w", newline="", encoding="utf-8") as fh: + writer = csv.DictWriter(fh, fieldnames=summary_fields, extrasaction="ignore") + writer.writeheader() + for row in results: + if "error" in row: + continue + flat = {k: row.get(k) for k in summary_fields} + # Rozkład sentymentów jako string JSON w komórce CSV + flat["sentiments_distribution"] = json.dumps( + row.get("sentiments_distribution", {}), ensure_ascii=False + ) + writer.writerow(flat) + + +def _summary_only(results: List[BenchmarkResult]) -> List[BenchmarkResult]: + """Zwraca wyniki bez pola per_text (lżejsza odpowiedź HTTP).""" + return [{k: v for k, v in r.items() if k != "per_text"} for r in results] + + +# --------------------------------------------------------------------------- +# Uruchomienie standalone +# --------------------------------------------------------------------------- + +if __name__ == "__main__": + import argparse + + parser = argparse.ArgumentParser(description="TruthScan NLP benchmark") + parser.add_argument( + "--adapters", default="roberta,xlm-roberta,norbert", + help="Przecinkowa lista adapterów (domyślnie: wszystkie)", + ) + parser.add_argument( + "--langs", default="en,pl,no", + help="Przecinkowa lista języków (domyślnie: en,pl,no)", + ) + parser.add_argument( + "--out-dir", default="benchmark_results", + help="Katalog wyjściowy dla plików JSON i CSV", + ) + args = parser.parse_args() + + adapter_names = [a.strip() for a in args.adapters.split(",")] + langs = [l.strip() for l in args.langs.split(",")] + out_dir = Path(args.out_dir) + + print(f"Uruchamiam benchmark: adaptery={adapter_names}, języki={langs}") + results = run_benchmark(adapter_names=adapter_names, langs=langs) + + json_path = out_dir / "benchmark.json" + csv_path = out_dir / "benchmark.csv" + export_json(results, json_path) + export_csv(results, csv_path) + + print(f"Wyniki zapisane: {json_path}, {csv_path}") + for r in _summary_only(results): + if "error" in r: + print(f" [{r['adapter_name']}] BŁĄD: {r['error']}") + else: + print( + f" [{r['adapter_name']:12s} / {r['language']}] " + f"avg={r['avg_inference_time_ms']} ms " + f"fake={r['avg_fake_probability']}% " + f"sentiments={r['sentiments_distribution']}" + ) diff --git a/TruthScan AI_backend/app/config.py b/TruthScan AI_backend/app/config.py new file mode 100644 index 0000000000000000000000000000000000000000..218f634f6a89c5bb594f6fa64bcf2db6f6d67c8b --- /dev/null +++ b/TruthScan AI_backend/app/config.py @@ -0,0 +1,50 @@ +""" +Centralna konfiguracja aplikacji oraz stałe wykorzystywane w wielu modułach. +""" + +import os +from pathlib import Path + +# Konfiguracja CORS +CORS_ALLOW_ORIGINS = ["*"] +CORS_ALLOW_CREDENTIALS = True +CORS_ALLOW_METHODS = ["*"] +CORS_ALLOW_HEADERS = ["*"] + +# Lista obsługiwanych źródeł RSS +NEWS_FEEDS = { + # Angielskie + "BBC": "https://feeds.bbci.co.uk/news/rss.xml", + "CNN": "http://rss.cnn.com/rss/edition.rss", + "NYTimes": "https://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml", + "Guardian": "https://www.theguardian.com/world/rss", + "AlJazeera": "https://www.aljazeera.com/xml/rss/all.xml", + # Polskie + "Money": "https://www.money.pl/rss/", + "PolsatNews": "https://www.polsatnews.pl/rss/wszystkie.xml", + "GazetaPrawna": "https://www.gazetaprawna.pl/rss.xml", + "SpidersWeb": "https://spidersweb.pl/feed", + "Bankier": "https://www.bankier.pl/rss/wiadomosci.xml", + # Norweskie + "NRK": "https://www.nrk.no/toppsaker.rss", + "VG": "https://www.vg.no/rss/feed/?limit=10", + "Dagbladet": "https://www.dagbladet.no/rss", + "Aftenposten": "https://www.aftenposten.no/rss", +} + +# Mapowanie wyników analizy sentymentu na etykiety językowe +SENTIMENT_MAP = { + "negative": {"pl": "Negatywne", "en": "Negative", "no": "Negativt"}, + "neutral": {"pl": "Neutralne", "en": "Neutral", "no": "Nøytralt"}, + "positive": {"pl": "Pozytywne", "en": "Positive", "no": "Positivt"}, +} + +# Ścieżka do pliku z zapisanymi artykułami +SAVED_FILE = Path("saved_articles.json") + +# Czas życia cache (sekundy) +CACHE_TTL_SECONDS = 120 + +# Konfiguracja Redis (jeśli używany jako backend cache) +REDIS_URL = os.getenv("REDIS_URL", "redis://localhost:6379") +CACHE_TTL = 300 diff --git a/TruthScan AI_backend/app/main.py b/TruthScan AI_backend/app/main.py new file mode 100644 index 0000000000000000000000000000000000000000..59a8a3f418099c9fe229569acec4a3fa423c1719 --- /dev/null +++ b/TruthScan AI_backend/app/main.py @@ -0,0 +1,57 @@ +""" +Główna konfiguracja i uruchomienie aplikacji FastAPI. +""" + +from fastapi import FastAPI +from fastapi.middleware.cors import CORSMiddleware +from fastapi_cache import FastAPICache +from fastapi_cache.backends.inmemory import InMemoryBackend + +# Konfiguracja CORS (źródła, metody, nagłówki itp.) +from .config import ( + CORS_ALLOW_ORIGINS, + CORS_ALLOW_CREDENTIALS, + CORS_ALLOW_METHODS, + CORS_ALLOW_HEADERS +) + +# Routery aplikacji +from .routes import misc, news, saved + +# Inicjalizacja aplikacji FastAPI +app = FastAPI() + +# Middleware CORS – umożliwia dostęp do API z innych domen +app.add_middleware( + CORSMiddleware, + allow_origins=CORS_ALLOW_ORIGINS, + allow_credentials=CORS_ALLOW_CREDENTIALS, + allow_methods=CORS_ALLOW_METHODS, + allow_headers=CORS_ALLOW_HEADERS, +) + +# Logika wykonywana przy starcie aplikacji +@app.on_event("startup") +async def startup(): + # Inicjalizacja cache w pamięci (np. do cache’owania newsów) + FastAPICache.init(InMemoryBackend(), prefix="news-cache") + print("✓ Cache initialized (5 minut TTL)") + +# Rejestracja routerów +app.include_router(misc.router) +app.include_router(news.router) +app.include_router(saved.router) + +# Endpoint główny – informacja o stanie API +@app.get("/") +async def root(): + return { + "message": "ThruScan API", + "status": "running", + "cached": True + } + +# Endpoint zdrowia – używany np. przez monitoring / load balancer +@app.get("/health") +async def health_check(): + return {"status": "healthy"} diff --git a/TruthScan AI_backend/app/models.py b/TruthScan AI_backend/app/models.py new file mode 100644 index 0000000000000000000000000000000000000000..dd3ce29d5ebc2b2c474b837717d899f03513809f --- /dev/null +++ b/TruthScan AI_backend/app/models.py @@ -0,0 +1,16 @@ +""" +Modele danych wykorzystywane do walidacji i serializacji artykułów. +""" + +from pydantic import BaseModel + + +class Article(BaseModel): + # Model artykułu wykorzystywany w komunikacji API + title: str + link: str + summary: str + published: str + sentiment: str + fake_probability: float + source: str diff --git a/TruthScan AI_backend/app/nlp_service.py b/TruthScan AI_backend/app/nlp_service.py new file mode 100644 index 0000000000000000000000000000000000000000..894250cb3a22c8184f454b7a66f49942f04acf21 --- /dev/null +++ b/TruthScan AI_backend/app/nlp_service.py @@ -0,0 +1,459 @@ +""" +Serwis NLP z architekturą plug-in. + +Każdy model jest reprezentowany przez adapter dziedziczący z ModelAdapter. +Publiczne API (analyze_news, analyze_news_batch) pozostaje niezmienione, +więc routes/news.py nie wymaga modyfikacji. +""" + +from abc import ABC, abstractmethod +from typing import Dict, Any, List, Optional +from concurrent.futures import ThreadPoolExecutor + +from transformers import pipeline + +from .config import SENTIMENT_MAP + + +# --------------------------------------------------------------------------- +# Klasa bazowa +# --------------------------------------------------------------------------- + +class ModelAdapter(ABC): + """ + Abstrakcyjny adapter modelu NLP. + + Każdy konkretny adapter musi zaimplementować: + - analyze_sentiment(text) -> {"label": str, "score": float} + - analyze_fake_news(text) -> {"labels": List[str], "scores": List[float]} + """ + + @property + @abstractmethod + def name(self) -> str: + """Unikalny identyfikator adaptera (np. 'roberta', 'xlm-roberta').""" + ... + + @property + @abstractmethod + def supported_languages(self) -> List[str]: + """Kody języków obsługiwanych przez model (np. ['pl', 'en', 'no']).""" + ... + + @abstractmethod + def analyze_sentiment(self, text: str) -> Dict[str, Any]: + """ + Analizuje sentyment tekstu. + + Zwraca: + {"label": str, "score": float} + gdzie label to jedna z wartości: 'positive' | 'negative' | 'neutral' + """ + ... + + @abstractmethod + def analyze_fake_news(self, text: str) -> Dict[str, Any]: + """ + Klasyfikuje tekst jako real/fake. + + Zwraca: + {"labels": List[str], "scores": List[float]} + """ + ... + + +# --------------------------------------------------------------------------- +# Adaptery +# --------------------------------------------------------------------------- + +def _normalize_sentiment_label(raw_label: str) -> str: + """ + Normalizuje etykietę sentymentu z modelu do jednego z trzech wariantów: + 'positive' | 'negative' | 'neutral'. + + Obsługuje różne konwencje nazewnictwa stosowane przez modele HuggingFace. + """ + label = (raw_label or "").strip().lower() + + _POSITIVE = {"positive", "pos", "label_2", "2", "very positive"} + _NEGATIVE = {"negative", "neg", "label_0", "0", "very negative"} + + if label in _POSITIVE or label.startswith("pos"): + return "positive" + if label in _NEGATIVE or label.startswith("neg"): + return "negative" + return "neutral" + + +class RoBERTaAdapter(ModelAdapter): + """ + Adapter dla modeli anglojęzycznych (domyślny, zachowuje obecne zachowanie): + - sentyment : cardiffnlp/twitter-roberta-base-sentiment-latest + - fake news : facebook/bart-large-mnli (zero-shot) + """ + + def __init__(self) -> None: + self._sentiment_pipe = None + self._fake_pipe = None + + def _load(self) -> None: + if self._sentiment_pipe is None: + self._sentiment_pipe = pipeline( + "text-classification", + model="cardiffnlp/twitter-roberta-base-sentiment-latest", + return_all_scores=False, + ) + if self._fake_pipe is None: + self._fake_pipe = pipeline( + "zero-shot-classification", + model="facebook/bart-large-mnli", + ) + + @property + def name(self) -> str: + return "roberta" + + @property + def supported_languages(self) -> List[str]: + return ["en"] + + def analyze_sentiment(self, text: str) -> Dict[str, Any]: + self._load() + result = self._sentiment_pipe(text)[0] + return { + "label": _normalize_sentiment_label(result.get("label", "")), + "score": float(result.get("score", 0.0)), + } + + def analyze_fake_news(self, text: str) -> Dict[str, Any]: + self._load() + result = self._fake_pipe(text, candidate_labels=["real", "fake"]) + return {"labels": result["labels"], "scores": result["scores"]} + + +class XLMRoBERTaAdapter(ModelAdapter): + """ + Adapter wielojęzyczny (pl, en, no): + - sentyment : cardiffnlp/twitter-xlm-roberta-base-sentiment + - fake news : facebook/bart-large-mnli (zero-shot, transfer między językami) + """ + + def __init__(self) -> None: + self._sentiment_pipe = None + self._fake_pipe = None + + def _load(self) -> None: + if self._sentiment_pipe is None: + self._sentiment_pipe = pipeline( + "text-classification", + model="cardiffnlp/twitter-xlm-roberta-base-sentiment", + return_all_scores=False, + ) + if self._fake_pipe is None: + self._fake_pipe = pipeline( + "zero-shot-classification", + model="facebook/bart-large-mnli", + ) + + @property + def name(self) -> str: + return "xlm-roberta" + + @property + def supported_languages(self) -> List[str]: + return ["pl", "en", "no"] + + def analyze_sentiment(self, text: str) -> Dict[str, Any]: + self._load() + result = self._sentiment_pipe(text)[0] + return { + "label": _normalize_sentiment_label(result.get("label", "")), + "score": float(result.get("score", 0.0)), + } + + def analyze_fake_news(self, text: str) -> Dict[str, Any]: + self._load() + result = self._fake_pipe(text, candidate_labels=["real", "fake"]) + return {"labels": result["labels"], "scores": result["scores"]} + + +class HerBERTAdapter(ModelAdapter): + """ + Adapter dla języka polskiego (HerBERT): + - sentyment : allegro/herbert-base-cased + UWAGA: to model bazowy – wymaga fine-tuningu na zbiorze sentymentu + (np. PolEmo 2.0). Aby podmienić checkpoint, przekaż sentiment_model + w konstruktorze lub zmień SENTIMENT_MODEL przed pierwszym użyciem. + - fake news : facebook/bart-large-mnli (zero-shot, transfer EN→PL) + """ + + SENTIMENT_MODEL: str = "allegro/herbert-base-cased" + + def __init__(self, sentiment_model: Optional[str] = None) -> None: + self._sentiment_model_id = sentiment_model or self.SENTIMENT_MODEL + self._sentiment_pipe = None + self._fake_pipe = None + + def _load(self) -> None: + if self._sentiment_pipe is None: + self._sentiment_pipe = pipeline( + "text-classification", + model=self._sentiment_model_id, + return_all_scores=False, + ) + if self._fake_pipe is None: + self._fake_pipe = pipeline( + "zero-shot-classification", + model="facebook/bart-large-mnli", + ) + + @property + def name(self) -> str: + return "herbert" + + @property + def supported_languages(self) -> List[str]: + return ["pl"] + + def analyze_sentiment(self, text: str) -> Dict[str, Any]: + self._load() + result = self._sentiment_pipe(text)[0] + return { + "label": _normalize_sentiment_label(result.get("label", "")), + "score": float(result.get("score", 0.0)), + } + + def analyze_fake_news(self, text: str) -> Dict[str, Any]: + self._load() + result = self._fake_pipe(text, candidate_labels=["real", "fake"]) + return {"labels": result["labels"], "scores": result["scores"]} + + +class NorBERTAdapter(ModelAdapter): + """ + Adapter dla języka norweskiego (NorBERT 3): + - sentyment : ltgoslo/norbert3-base + UWAGA: to model bazowy – wymaga fine-tuningu na zbiorze sentymentu + (np. NoReC). Aby podmienić checkpoint, przekaż sentiment_model + w konstruktorze lub zmień SENTIMENT_MODEL przed pierwszym użyciem. + - fake news : facebook/bart-large-mnli (zero-shot, transfer EN→NO) + """ + + SENTIMENT_MODEL: str = "ltgoslo/norbert3-base" + + def __init__(self, sentiment_model: Optional[str] = None) -> None: + self._sentiment_model_id = sentiment_model or self.SENTIMENT_MODEL + self._sentiment_pipe = None + self._fake_pipe = None + + def _load(self) -> None: + if self._sentiment_pipe is None: + self._sentiment_pipe = pipeline( + "text-classification", + model=self._sentiment_model_id, + return_all_scores=False, + ) + if self._fake_pipe is None: + self._fake_pipe = pipeline( + "zero-shot-classification", + model="facebook/bart-large-mnli", + ) + + @property + def name(self) -> str: + return "norbert" + + @property + def supported_languages(self) -> List[str]: + return ["no"] + + def analyze_sentiment(self, text: str) -> Dict[str, Any]: + self._load() + result = self._sentiment_pipe(text)[0] + return { + "label": _normalize_sentiment_label(result.get("label", "")), + "score": float(result.get("score", 0.0)), + } + + def analyze_fake_news(self, text: str) -> Dict[str, Any]: + self._load() + result = self._fake_pipe(text, candidate_labels=["real", "fake"]) + return {"labels": result["labels"], "scores": result["scores"]} + + +# --------------------------------------------------------------------------- +# Rejestr adapterów — lazy initialization +# --------------------------------------------------------------------------- + +# Przy imporcie pusty — żaden adapter nie jest tworzony ani ładowany. +# Adaptery są instancjonowane dopiero przy pierwszym wywołaniu get_adapter(). +_REGISTRY: Dict[str, ModelAdapter] = {} + +# Adapter aktywny globalnie; None oznacza „jeszcze nie wybrano". +_active_adapter: Optional[ModelAdapter] = None + +# Executor do równoległej analizy (ograniczenie obciążenia CPU) +executor = ThreadPoolExecutor(max_workers=3) + +# Zbiór nazw wbudowanych adapterów — służy do walidacji przed inicjalizacją +_BUILTIN_ADAPTERS = {"roberta", "xlm-roberta", "herbert", "norbert"} + + +def _init_registry() -> None: + """ + Tworzy wbudowane adaptery i wpisuje je do rejestru. + + Wywoływana leniwie przy pierwszym get_adapter() lub set_active_adapter(). + Kolejne wywołania są bezoperacyjne (idempotentna). + """ + if _REGISTRY: + return + _REGISTRY["roberta"] = RoBERTaAdapter() + _REGISTRY["xlm-roberta"] = XLMRoBERTaAdapter() + _REGISTRY["herbert"] = HerBERTAdapter() + _REGISTRY["norbert"] = NorBERTAdapter() + + +def get_adapter(name: str) -> ModelAdapter: + """ + Zwraca adapter o podanej nazwie. + + Przy pierwszym wywołaniu inicjalizuje rejestr (bez ładowania wag modeli). + Rzuca ValueError dla nieznanej nazwy. + """ + _init_registry() + if name not in _REGISTRY: + raise ValueError( + f"Nieznany adapter: '{name}'. Dostępne: {list(_REGISTRY.keys())}" + ) + return _REGISTRY[name] + + +def set_active_adapter(name: str) -> None: + """ + Ustawia aktywny adapter dla całej aplikacji. + + Przykład: + from app.nlp_service import set_active_adapter + set_active_adapter("xlm-roberta") + """ + global _active_adapter + _active_adapter = get_adapter(name) + + +def register_adapter(adapter: ModelAdapter) -> None: + """ + Rejestruje nowy adapter (np. fine-tuned checkpoint) pod jego nazwą. + Może być wywołane przed lub po _init_registry(). + + Przykład: + norbert_ft = NorBERTAdapter(sentiment_model="user/norbert3-norec") + register_adapter(norbert_ft) + set_active_adapter("norbert") + """ + _init_registry() + _REGISTRY[adapter.name] = adapter + + +def _get_active_adapter() -> ModelAdapter: + """ + Zwraca aktualnie aktywny adapter. + Jeśli nie ustawiono, domyślnie inicjalizuje i zwraca RoBERTaAdapter. + """ + global _active_adapter + if _active_adapter is None: + _active_adapter = get_adapter("roberta") + return _active_adapter + + +# --------------------------------------------------------------------------- +# Publiczne API – interfejs niezmieniony względem poprzedniej wersji +# --------------------------------------------------------------------------- + +def analyze_news( + text: str, + lang: str = "pl", + adapter: Optional[ModelAdapter] = None, +) -> dict: + """ + Analizuje pojedynczy tekst pod kątem sentymentu i fake news. + + Args: + text: Tekst do analizy. + lang: Kod języka wynikowych etykiet ('pl' | 'en' | 'no'). + adapter: Opcjonalny adapter; jeśli None, używa _active_adapter. + + Returns: + {"sentiment": str, "fake_probability": float, "sentiment_score": float} + """ + _neutral = SENTIMENT_MAP["neutral"].get(lang, "Neutral") + + if not text or len(text.strip()) < 10: + return {"sentiment": _neutral, "fake_probability": 0.0, "sentiment_score": 0.0} + + _adapter = adapter or _get_active_adapter() + + try: + sentiment_result = _adapter.analyze_sentiment(text) + fake_result = _adapter.analyze_fake_news(text) + except Exception: + return {"sentiment": _neutral, "fake_probability": 0.0, "sentiment_score": 0.0} + + label = sentiment_result.get("label", "neutral") + sentiment_translated = SENTIMENT_MAP.get(label, {}).get(lang, _neutral) + + fake_score = 0.0 + for lbl, score in zip(fake_result["labels"], fake_result["scores"]): + if lbl == "fake": + fake_score = score + break + + return { + "sentiment": sentiment_translated, + "fake_probability": round(fake_score * 100, 2), + "sentiment_score": round(float(sentiment_result.get("score", 0.0)), 2), + } + + +def analyze_news_batch( + texts: List[str], + lang: str = "pl", + adapter: Optional[ModelAdapter] = None, +) -> List[Dict[str, Any]]: + """ + Analiza wielu tekstów w trybie batch z użyciem ThreadPoolExecutor. + + Args: + texts: Lista tekstów do analizy. + lang: Kod języka wynikowych etykiet. + adapter: Opcjonalny adapter; jeśli None, używa _active_adapter. + """ + if not texts: + return [] + + _adapter = adapter or _get_active_adapter() + results: List[Dict[str, Any]] = [] + batch_size = 3 + + batches = [texts[i:i + batch_size] for i in range(0, len(texts), batch_size)] + + for batch in batches: + futures = [ + executor.submit(analyze_news, text, lang, _adapter) + for text in batch + ] + for future in futures: + try: + results.append(future.result()) + except Exception: + _neutral = SENTIMENT_MAP["neutral"].get(lang, "Neutral") + results.append( + {"sentiment": _neutral, "fake_probability": 0.0, "sentiment_score": 0.0} + ) + + return results + + +def analyze_news_single(text: str, lang: str) -> Dict[str, Any]: + """Wrapper dla analizy pojedynczego tekstu (używany w batch processing).""" + return analyze_news(text, lang) diff --git a/TruthScan AI_backend/app/routes/__init__.py b/TruthScan AI_backend/app/routes/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..e69de29bb2d1d6434b8b29ae775ad8c2e48c5391 diff --git a/TruthScan AI_backend/app/routes/misc.py b/TruthScan AI_backend/app/routes/misc.py new file mode 100644 index 0000000000000000000000000000000000000000..4a6f7e63981c0bd3e06dc8d9310e2e37d9d53f70 --- /dev/null +++ b/TruthScan AI_backend/app/routes/misc.py @@ -0,0 +1,68 @@ +""" +Endpointy związane z obsługą dostępnych źródeł wiadomości oraz narzędziami deweloperskimi. +""" + +from pathlib import Path +from typing import Optional + +from fastapi import APIRouter, Query + +from ..config import NEWS_FEEDS +from ..benchmark import run_benchmark, export_json, export_csv, _summary_only + +router = APIRouter() + +# Katalog, do którego endpoint zapisuje artefakty benchmarku +_BENCHMARK_OUT = Path("benchmark_results") + + +@router.get("/sources") +def get_sources(): + return list(NEWS_FEEDS.keys()) + + +@router.get("/benchmark") +def benchmark( + adapters: Optional[str] = Query( + default=None, + description="Przecinkowa lista adapterów: roberta,xlm-roberta,norbert", + ), + langs: Optional[str] = Query( + default=None, + description="Przecinkowa lista języków: en,pl,no", + ), + save: bool = Query( + default=True, + description="Czy zapisać wyniki do JSON i CSV w katalogu benchmark_results/", + ), + full: bool = Query( + default=False, + description="Czy zwrócić szczegółowe wyniki per_text (domyślnie tylko podsumowanie)", + ), +): + """ + Uruchamia benchmark NLP i zwraca wyniki. + + Czas odpowiedzi zależy od liczby adapterów i języków — może wynosić kilkadziesiąt sekund + przy pierwszym uruchomieniu (lazy-loading modeli HuggingFace). + + Przykłady: + - GET /benchmark + - GET /benchmark?adapters=roberta,xlm-roberta&langs=en,pl + - GET /benchmark?full=true&save=false + """ + adapter_names = [a.strip() for a in adapters.split(",")] if adapters else None + lang_list = [l.strip() for l in langs.split(",")] if langs else None + + results = run_benchmark(adapter_names=adapter_names, langs=lang_list) + + if save: + export_json(results, _BENCHMARK_OUT / "benchmark.json") + export_csv(results, _BENCHMARK_OUT / "benchmark.csv") + + return { + "results": results if full else _summary_only(results), + "saved": save, + "output_dir": str(_BENCHMARK_OUT) if save else None, + } + diff --git a/TruthScan AI_backend/app/routes/news.py b/TruthScan AI_backend/app/routes/news.py new file mode 100644 index 0000000000000000000000000000000000000000..88bdb0d894f45855ab4283abbb779253cec7bc41 --- /dev/null +++ b/TruthScan AI_backend/app/routes/news.py @@ -0,0 +1,234 @@ +""" +Endpointy związane z pobieraniem, analizą i strumieniowaniem wiadomości. +""" + +import asyncio +import json +from typing import List + +from fastapi import APIRouter, HTTPException +from fastapi.responses import StreamingResponse +from fastapi_cache.decorator import cache + +from ..config import CACHE_TTL, NEWS_FEEDS +from ..rss_utils import fetch_feed, clean_html, get_from_cache, set_to_cache +from ..nlp_service import analyze_news + +router = APIRouter() + + +@router.get("/news/{source}") +def get_news(source: str, lang: str = "pl"): + # Walidacja źródła + if source not in NEWS_FEEDS: + raise HTTPException(status_code=404, detail="Źródło nieobsługiwane") + + url = NEWS_FEEDS[source] + + # Pobranie RSS + try: + feed = fetch_feed(url) + except Exception as e: + raise HTTPException(status_code=500, detail=f"Błąd pobierania newsów: {str(e)}") + + if not feed.entries: + return {"source": source, "articles": []} + + MAX_ARTICLES = 5 + articles: List[dict] = [] + + # Przetwarzanie i analiza artykułów + for entry in feed.entries[:MAX_ARTICLES]: + summary = clean_html(entry.get("summary", "Brak opisu")) + text_to_analyze = (summary or "").strip() or entry.get("title", "") + analysis = analyze_news(text_to_analyze, lang) + + articles.append({ + "title": entry.get("title", "Bez tytułu"), + "link": entry.get("link", ""), + "summary": summary, + "published": entry.get("published", "Brak daty"), + "source": source, + **analysis + }) + + return {"source": source, "articles": articles} + + +# Strumieniowanie newsów przez SSE +@router.get("/stream-news/{source}") +async def stream_news(source: str, lang: str = "pl"): + if source not in NEWS_FEEDS: + raise HTTPException(status_code=404, detail="Źródło nieobsługiwane") + + async def event_generator(): + # Próba pobrania RSS z cache + try: + cached = get_from_cache(source) + if cached: + feed = cached + else: + feed = fetch_feed(NEWS_FEEDS[source]) + set_to_cache(source, feed) + except Exception as e: + yield f"event: backend_error\ndata: {json.dumps({'message': str(e)})}\n\n" + yield f"event: done\ndata: {json.dumps({'count': 0})}\n\n" + return + + entries = feed.entries or [] + MAX_ARTICLES = 5 + to_send = entries[:MAX_ARTICLES] + + # Metadane dla klienta + yield f"event: meta\ndata: {json.dumps({'total': len(to_send)})}\n\n" + + sent = 0 + for entry in to_send: + summary = clean_html(entry.get("summary", "Brak opisu")) + text_to_analyze = (summary or "").strip() or entry.get("title", "") + + # Analiza NLP uruchamiana w executorze (CPU-bound) + loop = asyncio.get_event_loop() + analysis = await loop.run_in_executor( + None, lambda: analyze_news(text_to_analyze, lang) + ) + + article = { + "title": entry.get("title", "Bez tytułu"), + "link": entry.get("link", ""), + "summary": summary, + "published": entry.get("published", "Brak daty"), + "source": source, + **analysis, + } + + yield f"data: {json.dumps(article, ensure_ascii=False)}\n\n" + sent += 1 + await asyncio.sleep(0.3) + + yield f"event: done\ndata: {json.dumps({'count': sent})}\n\n" + + return StreamingResponse(event_generator(), media_type="text/event-stream") + + +@router.get("/api/charts/summary") +@cache(expire=CACHE_TTL) +async def get_charts_summary(lang: str = "pl"): + """ + Szybkie statystyki zbiorcze dla wszystkich źródeł + (uproszczona analiza oparta na tytułach). + """ + sources = [ + "BBC", "CNN", "NYTimes", "Guardian", "AlJazeera", + "PolsatNews", "Money", "Bankier", "SpidersWeb", "GazetaPrawna" + ] + + summary_data = {} + + for source in sources: + if source not in NEWS_FEEDS: + continue + + try: + feed = fetch_feed(NEWS_FEEDS[source]) + + if not feed.entries: + summary_data[source] = { + "count": 0, + "emotions": {"Pozytywne": 0, "Negatywne": 0, "Neutralne": 0} + } + continue + + # Analiza tylko kilku tytułów (szybko) + articles = feed.entries[:3] + emotion_counts = {"Pozytywne": 0, "Negatywne": 0, "Neutralne": 0} + + for entry in articles: + title = entry.get("title", "").lower() + if any(word in title for word in ["good", "positive", "gain", "up", "success", "dobry", "wzrost", "zysk"]): + emotion_counts["Pozytywne"] += 1 + elif any(word in title for word in ["bad", "negative", "fall", "down", "loss", "crisis", "zły", "spadek", "kryzys"]): + emotion_counts["Negatywne"] += 1 + else: + emotion_counts["Neutralne"] += 1 + + summary_data[source] = { + "count": len(feed.entries), + "analyzed": len(articles), + "emotions": emotion_counts, + "latest_title": articles[0].get("title", "")[:50] if articles else "" + } + + except Exception: + summary_data[source] = { + "count": 0, + "emotions": {"Pozytywne": 0, "Negatywne": 0, "Neutralne": 0} + } + + return { + "summary": summary_data, + "total_sources": len(summary_data), + "cache_ttl": CACHE_TTL + } + + +@router.get("/emotion-stats/{source}") +@cache(expire=CACHE_TTL) +def get_emotion_stats(source: str, lang: str = "pl"): + # Statystyki emocji dla jednego źródła + if source not in NEWS_FEEDS: + raise HTTPException(status_code=404, detail="Źródło nieobsługiwane") + + try: + news_data = get_news(source, lang) + articles = news_data.get("articles", []) + except Exception as e: + raise HTTPException(status_code=500, detail=f"Błąd generowania statystyk: {str(e)}") + + emotion_counts = {"Pozytywne": 0, "Negatywne": 0, "Neutralne": 0} + + for article in articles: + sentiment = article.get("sentiment", "Neutralne") + if sentiment in emotion_counts: + emotion_counts[sentiment] += 1 + + total = len(articles) + emotion_percentages = { + emotion: (round((count / total) * 100, 2) if total > 0 else 0) + for emotion, count in emotion_counts.items() + } + + return { + "source": source, + "total_articles": total, + "emotion_counts": emotion_counts, + "emotion_percentages": emotion_percentages + } + + +@router.get("/charts-data") +@cache(expire=CACHE_TTL) +async def get_all_charts_data(lang: str = "pl"): + """ + Kompatybilność ze starszym frontendem – agreguje dane ze wszystkich źródeł. + """ + sources = [ + "BBC", "CNN", "NYTimes", "Guardian", "AlJazeera", + "PolsatNews", "Money", "Bankier", "SpidersWeb", "GazetaPrawna" + ] + + async def fetch_stats(source): + try: + return {source: await get_emotion_stats(source, lang)} + except Exception: + return {source: None} + + tasks = [fetch_stats(source) for source in sources] + results = await asyncio.gather(*tasks, return_exceptions=True) + + all_data = {} + for result in results: + if isinstance(result, dict): + all_data.update(result) + + return {"charts": all_data} diff --git a/TruthScan AI_backend/app/routes/saved.py b/TruthScan AI_backend/app/routes/saved.py new file mode 100644 index 0000000000000000000000000000000000000000..626f164028f5f50d023be04e8e25f18732ca6b97 --- /dev/null +++ b/TruthScan AI_backend/app/routes/saved.py @@ -0,0 +1,34 @@ +""" +Endpointy związane z zapisywaniem i zarządzaniem zapisanymi artykułami. +""" + +from fastapi import APIRouter, HTTPException, Request + +from ..storage import read_all, append_article, delete_by_title +from ..models import Article + +router = APIRouter() + + +@router.get("/saved-articles") +def get_saved_articles(): + # Zwraca listę wszystkich zapisanych artykułów + return read_all() + + +@router.post("/save-article") +def save_article(article: Article): + # Zapisuje nowy artykuł do magazynu danych + return append_article(article.dict()) + + +@router.delete("/delete-article") +async def delete_article(request: Request): + # Usuwa artykuł na podstawie tytułu przekazanego w body requestu + data = await request.json() + title = data.get("title") + + if not title: + raise HTTPException(status_code=400, detail="Brak pola 'title'") + + return delete_by_title(title) diff --git a/TruthScan AI_backend/app/rss_utils.py b/TruthScan AI_backend/app/rss_utils.py new file mode 100644 index 0000000000000000000000000000000000000000..3e9bcc7ee09e74198cc9cbfc3a057a306c2a2b9e --- /dev/null +++ b/TruthScan AI_backend/app/rss_utils.py @@ -0,0 +1,47 @@ +""" +Narzędzia pomocnicze do pobierania i przetwarzania kanałów RSS oraz cache w pamięci. +""" + +import time +import requests +import feedparser +from bs4 import BeautifulSoup +from typing import Dict, Any + +from .config import CACHE_TTL_SECONDS + +# Prosty cache w pamięci (key -> (value, timestamp)) +_cache_data: Dict[str, tuple[Any, float]] = {} + + +def clean_html(text: str) -> str: + # Usuwa znaczniki HTML z treści RSS + return BeautifulSoup(text or "", "html.parser").get_text() + + +def get_from_cache(key: str): + # Pobiera dane z cache, jeśli nie przekroczyły TTL + entry = _cache_data.get(key) + if not entry: + return None + + value, ts = entry + if time.time() - ts > CACHE_TTL_SECONDS: + del _cache_data[key] + return None + + return value + + +def set_to_cache(key: str, value): + # Zapisuje dane do cache wraz z timestampem + _cache_data[key] = (value, time.time()) + + +def fetch_feed(url: str): + # Pobiera i parsuje kanał RSS z ustawionym User-Agent + headers = {"User-Agent": "Mozilla/5.0 (compatible; ThruScanBot/1.0)"} + resp = requests.get(url, timeout=7, headers=headers) + resp.raise_for_status() + + return feedparser.parse(resp.content) diff --git a/TruthScan AI_backend/app/storage.py b/TruthScan AI_backend/app/storage.py new file mode 100644 index 0000000000000000000000000000000000000000..70dd2750865ee94fb2537fbf8f38c5f37bd29def --- /dev/null +++ b/TruthScan AI_backend/app/storage.py @@ -0,0 +1,60 @@ +""" +Warstwa dostępu do danych dla zapisanych artykułów (plik JSON). +""" + +import json +import threading +from fastapi import HTTPException + +from .config import SAVED_FILE + +# Blokada wątków dla operacji zapisu/odczytu +_lock = threading.Lock() + + +def ensure_file(): + # Tworzy plik danych, jeśli nie istnieje + if not SAVED_FILE.exists(): + SAVED_FILE.write_text("[]", encoding="utf-8") + + +def read_all(): + # Zwraca wszystkie zapisane artykuły + ensure_file() + try: + return json.loads(SAVED_FILE.read_text(encoding="utf-8")) + except json.JSONDecodeError: + return [] + + +def append_article(article: dict): + # Dodaje nowy artykuł do pliku JSON + ensure_file() + try: + with _lock, open(SAVED_FILE, "r+", encoding="utf-8") as f: + data = json.load(f) + data.append(article) + f.seek(0) + json.dump(data, f, ensure_ascii=False, indent=4) + + return {"message": "Artykuł zapisany pomyślnie."} + except Exception as e: + raise HTTPException(status_code=500, detail=str(e)) + + +def delete_by_title(title: str): + # Usuwa artykuł na podstawie tytułu + ensure_file() + try: + with _lock, open(SAVED_FILE, "r+", encoding="utf-8") as f: + saved = json.load(f) + new_saved = [ + a for a in saved if a.get("title") != title + ] + f.seek(0) + f.truncate() + json.dump(new_saved, f, ensure_ascii=False, indent=4) + + return {"message": "Artykuł usunięty."} + except Exception as e: + raise HTTPException(status_code=500, detail=str(e)) diff --git a/TruthScan AI_backend/requirements.txt b/TruthScan AI_backend/requirements.txt new file mode 100644 index 0000000000000000000000000000000000000000..55f91d0008251da32ed77e9156c8efae17d16805 --- /dev/null +++ b/TruthScan AI_backend/requirements.txt @@ -0,0 +1,22 @@ +# Zależności backendu aplikacji ThruScan + +# Framework API i serwer ASGI +fastapi==0.115.12 +uvicorn==0.34.3 +starlette==0.46.2 + +# Walidacja danych i modele +pydantic==2.11.7 +pydantic_core==2.33.2 +annotated-types==0.7.0 +typing-extensions>=4.12.2 + +# NLP i uczenie maszynowe +transformers==4.44.2 +tokenizers==0.19.1 +torch==2.3.1 + +# Przetwarzanie RSS i HTML +feedparser==6.0.11 +beautifulsoup4==4.12.3 +requests==2.31.0 diff --git a/TruthScan AI_backend/setup_test_env.bat b/TruthScan AI_backend/setup_test_env.bat new file mode 100644 index 0000000000000000000000000000000000000000..3078aac6859d609e1840778d3894e28865219024 --- /dev/null +++ b/TruthScan AI_backend/setup_test_env.bat @@ -0,0 +1,21 @@ +@echo off +echo ======================================== +echo 🛠 Przygotowanie środowiska testowego +echo ======================================== + +echo. +echo 📦 Instalowanie wymaganych bibliotek... +pip install requests pandas + +echo. +echo 🔍 Sprawdzanie czy backend działa... +timeout /t 3 /nobreak > nul + +echo. +echo 🚀 Uruchamianie testów... +python test_truthscan.py + +echo. +echo 📊 Testy zakończone! +echo Otwórz raport: truthscan_test_report.html +pause \ No newline at end of file diff --git a/TruthScan AI_backend/truthscan_report.html b/TruthScan AI_backend/truthscan_report.html new file mode 100644 index 0000000000000000000000000000000000000000..f3471edd8536dab5c849fe718d2169cbcd222a64 --- /dev/null +++ b/TruthScan AI_backend/truthscan_report.html @@ -0,0 +1,417 @@ + + +
+ +Artykułów: 5
+Fake: 25.5%
+Czas: 5.22s
+Artykułów: 5
+Fake: 17.2%
+Czas: 15.03s
+• Różnica w ryzyku dezinformacji: 8.3%
+• Różnica w czasie analizy: 9.81s
+• ✅ Spójna skuteczność detekcji między językami
+1. Who and what is in the Epstein files?...
2. David Walliams denies inappropriate behaviour after publisher drops him...
+1. Fałszywe oskarżenia na policji mogą zrujnować życie – sprawdź, kiedy grozi za ni...
2. Polacy za granicą: w tym kraju mieszka ich tylko dwoje. Zamiast polskiej zimy ma...
+| Test | +Status | +Szczegóły | +Czas [s] | +
|---|---|---|---|
| API Availability | +✅ PASS | +Swagger UI dostępny (200) | +2.07 | +
| GET /sources | +✅ PASS | +Znaleziono 10 źródeł | ✅ BBC dostępne | ✅ Gazeta Prawna dostępne | +2.06 | +
| GET /news/BBC | +✅ PASS | +Pobrano 5 artykułów | Źródło: BBC | Pola: title, sentiment, fake_probability, summary, link, published | Przykład: 'Who and what is in the Epstein files?' | +5.22 | +
| GET /news/GazetaPrawna | +✅ PASS | +Pobrano 5 artykułów | Źródło: GazetaPrawna | Pola: title, sentiment, fake_probability, summary, link, published | Przykład: 'Fałszywe oskarżenia na policji mogą zrujnować życi...' | +15.03 | +
| Source Comparison | +✅ PASS | +BBC: 5 art, 25.5% fake, 5.2s | Gazeta: 5 art, 17.2% fake, 15.0s | +- | +
| CRUD Operations - BBC | +✅ PASS | +Zapis: ✅ (2.07s) | Odczyt: ✅ (15 artykułów, 2.09s) | Usuwanie: ✅ (Status: 200) | +- | +
| CRUD Operations - GazetaPrawna | +✅ PASS | +Zapis: ✅ (2.06s) | Odczyt: ✅ (15 artykułów, 2.11s) | Usuwanie: ✅ (Status: 200) | +- | +
Artykułów: {self.comparison_data.get('source_count', {}).get('BBC', 0)}
+Fake: {fake_scores[0]:.1f}%
+Czas: {performance_times[0]:.2f}s
+Artykułów: {self.comparison_data.get('source_count', {}).get('Gazeta Prawna', 0)}
+Fake: {fake_scores[1]:.1f}%
+Czas: {performance_times[1]:.2f}s
+• Różnica w ryzyku dezinformacji: {abs(fake_scores[0] - fake_scores[1]):.1f}%
+• Różnica w czasie analizy: {abs(performance_times[0] - performance_times[1]):.2f}s
+ {"• ⚠️ Znacząca różnica w wykrywaniu dezinformacji między językami
" + if abs(fake_scores[0] - fake_scores[1]) > 10 else + "• ✅ Spójna skuteczność detekcji między językami
"} +{i+1}. {title}...
" + else: + html += "Brak przykładowych artykułów
" + + html += """ +{i+1}. {title}...
" + else: + html += "Brak przykładowych artykułów
" + + html += f""" +| Test | +Status | +Szczegóły | +Czas [s] | +
|---|---|---|---|
| {result['test_name']} | +{status_display} {result['status']} | +{result['details']} | +{duration} | +