ข้อมูลจากบทคัดย่อ
ยังไม่มีสรุปภาษาไทยสำหรับระเบียนนี้
ด้านล่างเป็นบทคัดย่อต้นฉบับภาษาอังกฤษจากข้อมูลบรรณานุกรม โปรดตรวจบทความต้นฉบับก่อนอ้างอิง
Safeguarding sensitive data in legal documents is essential for ensuring privacy, regulatory compliance, and secure information exchange. This paper introduces a legal text de-identification framework that targets five gaps in earlier research: low sample efficiency, insufficient focus on sensitive tokens, class imbalance, limited generalizability, and difficult hyperparameter tuning. Our approach employs a dual-agent architecture based on an enhanced trust region policy optimization (TRPO) algorithm. The enhanced TRPO adds an entropy-based regularization (ER) term to the advantage function. It also aggregates outcomes across multiple steps to improve stability and policy robustness. The first agent serves as an active learning (AL) module that selectively identifies the most informative unlabeled samples, reducing reliance on extensive labeled corpora. These prioritized samples are then processed by a second agent responsible for anonymization. The anonymization pipeline integrates local interpretable model-agnostic explanations (LIME) to evaluate and rank token-level importance. Complementary reward mechanisms address class imbalance and strengthen recognition of underrepresented entities. To enhance semantic representation, word embeddings are derived from a generative pre-trained transformer (GPT) model. To enrich the training data, a relational generative adversarial network framework, enhanced with gradient signal exclusion (RelGANGE), is used for online data augmentation. To increase sample diversity, the generator in RelGANGE ignores gradients from samples that receive high discriminator weights. Hyperparameter tuning is performed using a Homotopy-based optimization strategy to refine system performance. Evaluation of the Tribunal Arbitral du Basket-ball (TAB) and European Court of Human Rights (ECtHR) benchmark datasets yields F-measures of 92.414% and 92.637% across the reported settings. These results demonstrate practical value. The combined use of ER-TRPO, AL, LIME, GPT, RelGANGE, and Homotopy enables effective legal text de-identification.
เหตุผลที่อยู่ในฐานติดตาม
ระเบียนนี้ได้รับ Impact Signal 86/100 จากความใหม่ แหล่งเผยแพร่ ความร่วมมือ และสัญญาณในข้อมูลบรรณานุกรม คะแนนนี้ใช้จัดลำดับการติดตาม ไม่ใช่การตัดสินคุณภาพงานวิจัย
ประเด็นที่เกี่ยวข้อง: Topic Modeling · Authorship Attribution and Profiling · Hate Speech and Cyberbullying Detection
บทบาทของนักวิจัยและสถาบันไทย
Roohallah Alizadehsani · Chulalongkorn University
ข้อจำกัดของข้อมูล
หน้านี้เป็นระเบียนบรรณานุกรมและข้อมูลจากบทคัดย่อ ยังไม่ใช่บทวิเคราะห์ฉบับเต็มหรือการประเมินคุณภาพงานวิจัย ควรตรวจสอบ DOI และเอกสารต้นฉบับก่อนนำไปอ้างอิง