← People
Pedro Ortiz Suarez
interacted with
Common Crawl Foundation
met · ACL 2026 · 2026-07-07
Google Scholar
Papers in the feed →
Papers · 40
Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains
arXiv.org
2026-02-27
alphaXiv
arXiv
S2
How Should We Model the Probability of a Language?
VarDial@EAC
2026-02-09
alphaXiv
arXiv
S2
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Annual Meeting of the Association for Computational Linguistics
2026-01-25
alphaXiv
arXiv
S2
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
arXiv.org
2025-12-16
alphaXiv
arXiv
S2
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
arXiv.org
2025-12-12
alphaXiv
arXiv
S2
A Framework for Ethical Data Removal from Language Resources: An Example from Low-Resource Language Communities
Proceedings of the AAAI Symposium Series
2025-11-23
S2
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
Proceedings of the 4th Table Representation Learning Workshop
2025-06-30
alphaXiv
arXiv
S2
KréyoLID From Language Identification Towards Language Mining
arXiv.org
2025-03-09
alphaXiv
arXiv
S2
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem
Conference on Empirical Methods in Natural Language Processing
2025
S2
Building Data Infrastructure for Low-Resource Languages
Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2025)
2025
S2
Data Processing for the OpenGPT-X Model Family
arXiv.org
2024-10-11
alphaXiv
arXiv
S2
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
European Conference on Artificial Intelligence
2024-09-30
alphaXiv
arXiv
S2
Documenting Geographically and Contextually Diverse Language Data Sources
NEJLT
2024-09-12
S2
Molyé: A Corpus-based Approach to Language Contact in Colonial France
NLP4DH
2024-08-08
alphaXiv
arXiv
S2
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
Annual Meeting of the Association for Computational Linguistics
2024-06-13
alphaXiv
arXiv
S2
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages
International Conference on Language Resources and Evaluation
2024-05-01
S2
Tokenizer Choice For LLM Training: Negligible or Crucial?
NAACL-HLT
2023-10-12
alphaXiv
arXiv
S2
Semi-automatic staging area for high-quality structured data extraction from scientific literature
Science and Technology of Advanced Materials: Methods
2023-09-19
alphaXiv
arXiv
S2
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Neural Information Processing Systems
2023-03-07
alphaXiv
arXiv
S2
Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data
arXiv.org
2022-12-20
alphaXiv
arXiv
S2
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
arXiv.org
2022-11-09
alphaXiv
arXiv
S2
Automatic extraction of materials and properties from superconductors scientific literature
Science and Technology of Advanced Materials: Methods
2022-10-26
alphaXiv
arXiv
S2
BERTrade: Using Contextual Embeddings to Parse Old French
International Conference on Language Resources and Evaluation
2022-06-01
S2
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French
International Conference on Language Resources and Evaluation
2022-02-18
alphaXiv
arXiv
S2
Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources
arXiv.org
2022-01-25
alphaXiv
arXiv
S2
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
International Conference on Language Resources and Evaluation
2022-01-17
alphaXiv
arXiv
S2
Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus
2021-06-23
S2
A dataset for automatic detection of places in (early) modern French texts
2021-05-28
S2
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Transactions of the Association for Computational Linguistics
2021-03-22
alphaXiv
arXiv
S2
SinNer@Clef-Hipe2020 : Sinful adaptation of SotA models for Named Entity Recognition in French and German
Conference and Labs of the Evaluation Forum
2020-09-23
S2
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
Annual Meeting of the Association for Computational Linguistics
2020-07-05
S2
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
Annual Meeting of the Association for Computational Linguistics
2020-06-11
alphaXiv
arXiv
S2
Les modèles de langue contextuels Camembert pour le français : impact de la taille et de l’hétérogénéité des données d’entrainement (C AMEM BERT Contextual Language Models for French: Impact of Training Data Size and Heterogeneity )
JEPTALNRECITAL
2020-06-08
S2
French Contextualized Word-Embeddings with a sip of CaBeRnet: a New French Balanced Reference Corpus
CMLC
2020-05-16
S2
Establishing a New State-of-the-Art for French Named Entity Recognition
International Conference on Language Resources and Evaluation
2020-05-11
alphaXiv
arXiv
S2
CamemBERT: a Tasty French Language Model
Annual Meeting of the Association for Computational Linguistics
2019-11-10
alphaXiv
arXiv
S2
How OCR Performance can Impact on the Automatic Extraction of Dictionary Content Structures
2019-09-18
S2
Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
2019-07-22
S2
ElChat: Adapting Chat Language Models Using Only Target Unlabeled Language Data
S2
A Data-driven Approach to Natural Language Processing for Contemporary and Historical French
S2