Hi, welcome to my Resume page. Please see my English Resume or my Chinese Resume for the printout version.
Experiences
Research Projects
- Agentic Manga Translation and Evaluation University of Oxford 2026--2026
- Developing a closed-loop agentic manga translation system with automatic evaluation for full-page visual-text translation quality.
- Built translation support for full-page content, including dialogue, narration, footnotes, and onomatopoeia.
- Designed an automatic evaluation system that detects both translation errors and text-artwork integration issues.
- Audio-based Heart Rate Monitoring University of Oxford/Marley Health 2023--2025
- Built a noise-robust audio-based resting heart-rate monitoring algorithm for pet wearables.
- Developed a noise-gating algorithm that improved robustness in audio-based animal resting heart-rate monitoring.
- Reduced p95 heart-rate error - prediction availability AUC from 19.05 to 7.03.
- Achieved 11.08 BPM p95 error and 3.25 BPM mean error at resting heart rate.
- Pipeline-style LLM-based Manga Translation Amazon JP. GK 2025--2026
- Built a pipeline-style LLM-based manga translation system with automatic quality evaluation.
- Developed a preserve-then-select technique that reduced content localization errors by 13% and translation errors by 45%.
- Developed an LLM-based automatic evaluation system with 0.87 correlation to human MQM judgments.
- LLM-informed Syntax Parsing University of Tokyo 2023--2025
- Developed LLM-informed methods for unsupervised constituency and dependency parsing by combining paraphrase-based resampling, reinforcement learning, and grammatical priors.
- Designed a paraphrase-based resampling method that mitigated spurious textual patterns and improved unsupervised constituency parsing accuracy by 8 absolute points across four languages.
- Estimated word--word mutual information with grammatical constraints for LLM-based dependency parsing, improving parsing accuracy by over 5%.
- Published the resulting work at ICLR 2025 Spotlight, ACL 2024, and NAACL 2024.
- Personal Information Identification using Sparse-AutoEncoder University of Tokyo 2023--2025
- Built a personal information identification model using LLM Sparse Autoencoder (SAE) features.
- Improved PII detection F1 from 72% to 87% by integrating SAE features with an LSTM decoder.
- Syntax-informed Semantic Dependency Parsing with Mixture-of-Experts University of Tokyo 2021--2022
- Developed a mixture-of-experts approach that conditions semantic dependency labels on automatically discovered syntactic patterns.
- Improved semantic dependency parsing accuracy by $\sim$1 absolute F1 point over state-of-the-art methods.
- Published the resulting work as an ACL 2022 Oral paper.
Last publications
-
ICLR Spotlight
Is language modeling sufficient for accurate unsupervised parsing? No, sentence-level semantic information greatly contribute to robust and accurate parsing.
View paper -
AAAI Oral
Proper prosody modeling helps with speech intelligibility.
View paper -
Findings of ACL
Does neural similarity capture substring-level semantic similarity? Not necessarily, substring-frequency among paraphrases might be a better choice.
View paper -
NAACL
A better LM-based mutual information estimate helps with dependency parsing, though the MI estimate often ignore import syntactic information.
View paper -
Journal of Natural Language Processing
This comment investigates the correlation between syntactic and semantic dependencies in semantic role labeling.
View paper -
Modeling Syntactic-Semantic Dependency Correlations in Semantic Role Labeling Using Mixture Models
May 2022ACL Oral
How will syntax better help semantic parsing? Just separately model the semantic dependency per syntactic pattern and cluster the pattern using variational inference.
View paper -
An Improved StarGAN for Emotional Voice Conversion: Enhancing Voice Quality and Data Augmentation
Sep 2021Interspeech
Presents an improved StarGAN architecture for emotional voice conversion. Two stage training helps with StarGAN generation quality.
View paper -
NLP-COVID Workshop @ EMNLP
This paper describes a large-scale system for aggregating worldwide information about the COVID-19 pandemic.
View paper -
Frontiers in Artificial Intelligence
This journal article details a pattern-based approach for named entity recognition in Chinese medical imaging reports.
-
A Bibliometric Analysis of the Research Status of the Technology Enhanced Language Learning
Sep 2018SETE@ICWL
This paper presents a bibliometric analysis of the research landscape in Technology Enhanced Language Learning (TELL).
Skills
ML/LLM
LLM Application, Multimodal Translation, Automated Evaluation, MLOps,Unsupervised Syntax Parsing, Audio-based Biosignals Analysis
Programming
Python, TypeScript, Bash
Frameworks
PyTorch, Lightning, vLLM, Hydra, TensorFlow, Hugging Face Accelerate, Megatron
MLOps
MLFlow, DeepEval
Infrastructure
AWS, Slurm, Cloudflare
Languages
Chinese (Native), Japanese (N1, Business), English (C1, Business Level)
Education
Grants and Awards
- 2023
DC2 Fellowship
Japan Society for the Promotion of Science
- 2024
Special Allowance for Outstanding Student
Japan Society for the Promotion of Science
- 2025
Travel Grant ($2000)
Association for the Advancement of Artificial Intelligence
- 2022
IST-RA Fellowship
The University of Tokyo