About VerusCite

Our Methodology

VerusCite is designed to provide rigorous verification of academic and legal citations. Our process involves several sophisticated steps to ensure accuracy:

  • Upload limits: Papers may be up to 1000 pages (and up to 100 MB per file).
  • Extraction: We use LLM agents to parse document text and extract structured citation data.
  • Verification: Extracted citations are verified using Crossref and web search. This allows us to verify all regularly cited sources, like peer review articles, books, web pages, and more.
  • Reasoning: Our system employs advanced LLM reasoning to compare cited metadata against found results, identifying potential hallucinations or mismatches.

In our public benchmark of 36 documents and over 2,000 citations, VerusCite has an under 1% false positive rate for hallucinations, and an over 70% recall rate for citations with some type of error.

Data Privacy & Retention

We prioritize your privacy and data security:

  • Selective LLM Processing: Only identified sections with references are uploaded to the LLM; the main text of the document is not sent to any external LLM API.
  • No Long-Term API Retention: Our LLM providers are all zero-data retention, and our agents are configured with no long-term data retention for processed text.
  • PDF Deletion: Your original PDF files are deleted from our server immediately after the extraction process is complete.
  • Metadata Storage: We only retain the extracted citation metadata and verification results to allow you to review and download them.

Site analytics

We use a simple first-party pageview counter so we can understand which pages are used and roughly where visitors come from. This is not Google Analytics or any third-party ad tracker.

  • No marketing cookies: We do not set tracking cookies or use third-party analytics scripts for this purpose, so we do not show a cookie consent banner for site metrics.
  • What we record: page path, external referrer (if any), approximate location derived from IP (city/region/country when available), a coarse browser user-agent string, and a daily-rotating hash used only to estimate unique visitors.
  • What we do not store: raw IP addresses, and we do not sell or share this data with advertisers.
  • Where it lives: counts are stored in our own database and are only visible to site administrators—not to other users of VerusCite.

Approximate location data is provided by DB-IP (free City Lite database). IP geolocation is inherently imprecise (VPNs and shared networks can map to the wrong city).

About the Developer

VerusCite was developed by Andrew P. Wheeler, PhD, a data scientist and former professor of criminology at the University of Texas at Dallas. Andrew has published extensively in peer-reviewed journals and serves on the editorial boards of several academic journals. His experience as an author, reviewer, and editor gives him firsthand knowledge of the work involved in publishing and reviewing scholarly research—and of how important accurate citations are to that process.

Dr. Wheeler has expertise in building cost-effective LLM applications. See his book, Large Language Models for Mortals: A Practical Guide for Analysts with Python, for examples of his expertise in document processing and agentic applications.

Learn more about Andrew's research and writing at andrewpwheeler.com, or about his data science consulting and training work at CRIME De-Coder.