Do Text Anonymizers Generalize Across Contexts? Extending RAT-Bench to Malaysian Microdata and PII
Abstract
Text anonymizers are often used before sensitive text enters foundation-model workflows, including training, fine-tuning, retrieval, and user-facing inference. However, these tools are commonly evaluated on English text with U.S.-centric identifiers and cultural context. We extend RAT-Bench, an existing primarily English-language re-identification benchmark, to the Malaysian context using local microdata, Malaysian personally identifiable information (PII), and culturally grounded transcripts in Malaysian English and Bahasa Malaysia. We evaluate NER- and LLM-based anonymizers with an LLM attacker that infers attributes from anonymized text, measuring both re-identification success and text utility. Across Malaysian English and Bahasa Malaysia, LLM anonymizers provide the strongest explicit privacy-utility trade-offs, reducing Easy/Hard re-identification risk to 23-29% while preserving BLEU scores of 0.77-0.94. These results show that anonymization benchmarks are not context-neutral: local identifiers, language use, and cultural context should be considered before deploying anonymization tools in target foundation-model pipelines.