Skip to content
OpenTrain AIFor AI Companies

GlotLID: Language Identification for Low-Resource Languages

Amir Kargaran, Ayyoob Imani, François Yvon, Hinrich SchuetzePublished Jan 1, 2023
DOI Publisher
Researcher verdict
Context only
Use as context only
Benchmark evidence
Missing
Not verified yet
Time to first repro
A few days
Plan setup time
Risk flags
2
Review before use

Abstract

Domain fit: AI-adjacent · Paper appears method- or tooling-adjacent to AI workflows with partial ecosystem coverage.

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages.However, there is no LID available that (i) covers a wide range of low-resource languages, (ii) is rigorously evaluated and reliable and (iii) efficient and easy to use.Here, we publish GlotLID-M, an LID model that satisfies the desiderata of wide coverage, reliability and efficiency.It identifies 1665 languages, a large increase in coverage compared to prior work.In our experiments, GlotLID-M outperforms four baselines (CLD3, FT176, OpenLID and NLLB) when balancing F1 and false positive rate (FPR).We analyze the unique challenges that low-resource LID poses: incorrect corpus metadata, leakage from high-resource languages, difficulty separating closely related languages, handling of macrolanguage vs varieties and in general noisy data.We hope that integrating GlotLID-M into dataset creation pipelines will improve quality and enhance accessibility of NLP technology for low-resource languages and cultures.GlotLID-M model, code, and list of data sources are available: https: //github.com/cisnlp/GlotLID.

Results and benchmarks

Freshness tier: cold
Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages.However, there is no LID available that (i) covers a wide range of low-resource languages, (ii) is rigorously evaluated and reliable and (iii) efficient and easy to use.Here, we publish GlotLID-M, an LID model that satisfies the desiderata of wide coverage, reliability and efficiency.It identifies 1665 languages, a large increase in coverage compared to prior work.In our experiments, GlotLID-M outperforms four baselines (CLD3, FT176, OpenLID and NLLB) when balancing F1 and false positive rate (FPR).We analyze the unique challenges that low-resource LID poses: incorrect corpus metadata, leakage from high-resource languages, difficulty separating closely related languages, handling of macrolanguage vs varieties and in general noisy data.We hope that integrating GlotLID-M into dataset creation pipelines will improve quality and enhance accessibility of NLP technology for low-resource languages and cultures.GlotLID-M model, code, and list of data sources are available: https: //github.com/cisnlp/GlotLID.

Implementation

No direct implementation yet

Maintained implementation evidence is not confirmed for this paper yet.

Use the implementation status and reproduction sections for the current action plan.

Implementation evidence summary
Confidence: low

Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.

Reproduction risks
  • Estimate is based on paper-only reproduction flow

Reproduction readiness

Time to first repro: days
Last checked: Aug 20, 2026

No repo

No verified implementation available

  • No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.

Hardware requirements

  • Expect multi-day setup/compute for meaningful reproduction based on current guidance.

Hugging Face artifacts

No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.

Tip: start with models, then check datasets and spaces if you need evaluation data or demos.

Research context

15

Citations

0

References

Tasks

Computer science, Metadata, Publication, Resource (disambiguation), Identification (biology), Reliability (semiconductor), World Wide Web

Methods

Information retrieval

Domains

Natural language processing, Artificial intelligence

Evaluation and human feedback data

Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.

Open in HFEPX
Explore similar papers