← People
Joël Niklaus
following
University of Bern
@joelniklaus
Papers in the feed →
Papers · 9
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
arXiv.org
2026-04-15
alphaXiv
arXiv
S2
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Annual Meeting of the Association for Computational Linguistics
2025-08-06
alphaXiv
arXiv
S2
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
arXiv.org
2025-05-19
alphaXiv
arXiv
S2
SwiLTra-Bench: The Swiss Legal Translation Benchmark
Annual Meeting of the Association for Computational Linguistics
2025-03-03
alphaXiv
arXiv
S2
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords
Annual Meeting of the Association for Computational Linguistics
2025
S2
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
International Conference on Learning Representations
2024-11-29
alphaXiv
arXiv
S2
Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland
Conference on Empirical Methods in Natural Language Processing
2024-10-17
alphaXiv
arXiv
S2
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
Annual Meeting of the Association for Computational Linguistics
2024-10-17
alphaXiv
arXiv
S2
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
Neural Information Processing Systems
2023
S2