← People
Pedro Ortiz Suarez

Pedro Ortiz Suarez interacted with

Common Crawl Foundation
met · ACL 2026 · 2026-07-07
Google ScholarPapers in the feed →

Papers · 40
  1. Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains
    arXiv.org2026-02-27alphaXiv arXiv S2
  2. How Should We Model the Probability of a Language?
    VarDial@EAC2026-02-09alphaXiv arXiv S2
  3. CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
    Annual Meeting of the Association for Computational Linguistics2026-01-25alphaXiv arXiv S2
  4. Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
    arXiv.org2025-12-16alphaXiv arXiv S2
  5. SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
    arXiv.org2025-12-12alphaXiv arXiv S2
  6. A Framework for Ethical Data Removal from Language Resources: An Example from Low-Resource Language Communities
    Proceedings of the AAAI Symposium Series2025-11-23S2
  7. Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
    Proceedings of the 4th Table Representation Learning Workshop2025-06-30alphaXiv arXiv S2
  8. KréyoLID From Language Identification Towards Language Mining
    arXiv.org2025-03-09alphaXiv arXiv S2
  9. Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem
    Conference on Empirical Methods in Natural Language Processing2025S2
  10. Building Data Infrastructure for Low-Resource Languages
    Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2025)2025S2
  11. Data Processing for the OpenGPT-X Model Family
    arXiv.org2024-10-11alphaXiv arXiv S2
  12. Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
    European Conference on Artificial Intelligence2024-09-30alphaXiv arXiv S2
  13. Documenting Geographically and Contextually Diverse Language Data Sources
    NEJLT2024-09-12S2
  14. Molyé: A Corpus-based Approach to Language Contact in Colonial France
    NLP4DH2024-08-08alphaXiv arXiv S2
  15. mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
    Annual Meeting of the Association for Computational Linguistics2024-06-13alphaXiv arXiv S2
  16. A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages
    International Conference on Language Resources and Evaluation2024-05-01S2
  17. Tokenizer Choice For LLM Training: Negligible or Crucial?
    NAACL-HLT2023-10-12alphaXiv arXiv S2
  18. Semi-automatic staging area for high-quality structured data extraction from scientific literature
    Science and Technology of Advanced Materials: Methods2023-09-19alphaXiv arXiv S2
  19. The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
    Neural Information Processing Systems2023-03-07alphaXiv arXiv S2
  20. Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data
    arXiv.org2022-12-20alphaXiv arXiv S2
  21. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
    arXiv.org2022-11-09alphaXiv arXiv S2
  22. Automatic extraction of materials and properties from superconductors scientific literature
    Science and Technology of Advanced Materials: Methods2022-10-26alphaXiv arXiv S2
  23. BERTrade: Using Contextual Embeddings to Parse Old French
    International Conference on Language Resources and Evaluation2022-06-01S2
  24. From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French
    International Conference on Language Resources and Evaluation2022-02-18alphaXiv arXiv S2
  25. Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources
    arXiv.org2022-01-25alphaXiv arXiv S2
  26. Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
    International Conference on Language Resources and Evaluation2022-01-17alphaXiv arXiv S2
  27. Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus
    2021-06-23S2
  28. A dataset for automatic detection of places in (early) modern French texts
    2021-05-28S2
  29. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
    Transactions of the Association for Computational Linguistics2021-03-22alphaXiv arXiv S2
  30. SinNer@Clef-Hipe2020 : Sinful adaptation of SotA models for Named Entity Recognition in French and German
    Conference and Labs of the Evaluation Forum2020-09-23S2
  31. Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
    Annual Meeting of the Association for Computational Linguistics2020-07-05S2
  32. A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
    Annual Meeting of the Association for Computational Linguistics2020-06-11alphaXiv arXiv S2
  33. Les modèles de langue contextuels Camembert pour le français : impact de la taille et de l’hétérogénéité des données d’entrainement (C AMEM BERT Contextual Language Models for French: Impact of Training Data Size and Heterogeneity )
    JEPTALNRECITAL2020-06-08S2
  34. French Contextualized Word-Embeddings with a sip of CaBeRnet: a New French Balanced Reference Corpus
    CMLC2020-05-16S2
  35. Establishing a New State-of-the-Art for French Named Entity Recognition
    International Conference on Language Resources and Evaluation2020-05-11alphaXiv arXiv S2
  36. CamemBERT: a Tasty French Language Model
    Annual Meeting of the Association for Computational Linguistics2019-11-10alphaXiv arXiv S2
  37. How OCR Performance can Impact on the Automatic Extraction of Dictionary Content Structures
    2019-09-18S2
  38. Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
    2019-07-22S2
  39. ElChat: Adapting Chat Language Models Using Only Target Unlabeled Language Data
    S2
  40. A Data-driven Approach to Natural Language Processing for Contemporary and Historical French
    S2