
About ResLit
ResLit is an automatically curated database of antimicrobial resistance (AMR) genes and mutations, built by systematically mining the peer-reviewed scientific literature using a large language model pipeline.
Our Mission
ResLit serves two audiences with a shared need: access to the AMR literature at scale.
For curators working on databases like CARD, ResFinder, or institutional collections, ResLit acts as a discovery tool, surfacing genes, mutations, and papers that warrant manual review and curation, prioritized by evidence strength. Instead of searching the literature from scratch, curators can use ResLit to identify what is already well-supported, what is newly reported, and where their attention will have the most impact.
For scientists, ResLit provides a real-time snapshot of what the literature contains, both curated and uncurated. Researchers can explore resistance findings across organisms, antibiotics, and geographic regions, see how well-established each entry is, and follow every finding back to its source publication. Whether a gene is Confirmed across multiple reference databases or a Candidate reported in a single paper, that distinction is always visible.
How It Works
Literature Search
Over 1.5 million AMR-relevant publications were retrieved from PubMed using targeted search queries covering resistance genes, mutations, organisms, and antibiotic classes. 150,000 full-text articles were obtained through PubMed Central Open Access, publisher APIs, and institutional access agreements.
LLM-Based Extraction
Each paper is processed by a two-pass pipeline powered by Qwen3-30B-A3B, an open-source large language model. The first pass performs high-recall extraction of genes, mutations, organisms, antibiotics, and geographic locations. The second pass audits each extracted entity individually against the source text, verifying it reflects the paper's own experimental findings and not background citations or prior literature. Papers exceeding the model's context window are split into overlapping chunks and merged after extraction.
Normalization
Raw extracted entries undergo multi-layer normalization before entering the database. Gene names are resolved to canonical gene family names using a curated allele-to-family mapping, handling allele variants such as blaTEM-1, tetA(1), and aac(6')-Ib. Antibiotic names and abbreviations are standardized against a reference list covering all major drug classes. Mutation notation is normalized to standard nucleotide and protein change formats, with special handling for rRNA genes, promoter mutations, indels, and frameshifts.
Cross-Database Validation
Normalized entries are cross-referenced against three established reference databases: CARD, ResFinder, and the NCBI Reference Gene Catalog, using gene name as the primary matching key. This cross-referencing both validates pipeline precision and anchors ResLit findings within the broader AMR knowledge ecosystem.
Evidence Grading
Each entry is assigned one of four confidence tiers — Confirmed, Established, Supported, or Candidate — based on the number of external databases it appears in and the number of independent publications reporting it. This tiering gives users a transparent, reproducible signal of evidence strength without requiring manual review of every entry.
Evidence Tiers
Present in 2+ databases (CARD, ResFinder, NCBI RGC, ResLit) — independently curated by more than one reference source.
Present in exactly one external database — high confidence, but not yet cross-validated across references.
Extracted by ResLit from 3 or more independent publications — replicated in the literature but absent from external databases.
Extracted by ResLit from fewer than 3 papers — single or limited reports; highest false-positive risk and highest priority for manual review.
Data Coverage
ResLit covers resistance mechanisms across all major antibiotic classes, including beta-lactams, aminoglycosides, fluoroquinolones, glycopeptides, macrolides, tetracyclines, and polymyxins. Both acquired resistance genes and chromosomal mutations are represented, spanning a wide range of clinically and epidemiologically relevant bacterial species.
How to Cite
A manuscript describing ResLit is currently in preparation. Please check back for citation details.
Start Exploring
Browse our database of AMR genes and mutations, or learn how you can contribute as a curator.