I am a Ph.D. Candidate from the Department of Computer Science and Technology, Tsinghua University under the supervision of Prof. Xiaolin Hu. I am a member of the Tsinghua Statistical Artificial Intelligence & Learning (TSAIL) Group. My research interests include the application of artificial intelligence and machine learning in finance, insurance, speech, and audio processing. I am also the co-foudner of Promptlaw, a legal AI startup aimed at making legal services more accessible and affordable with the power of AI.
Prior to my PhD, I had over 8 years of industrial experience in data science, actuarial, and managerial roles in multinational corporations across Australia and the surrounding regions.
Ph.D. in Artificial Intelligence
Tsinghua University
MEng in Computer Science
University of New South Wales
BSc in Mathematics and Statistics
University of New South Wales

Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applications their usefulness depends on rare adverse scenarios rather than typical samples. Standard generative objectives prioritize bulk distributional fidelity, leaving low-probability tails vulnerable to localized optimization noise and making tail-dependent functionals unstable under finite simulation budgets. We introduce Diachronic Sample Integration (DSI), a test-time inference framework that ensembles generated samples across checkpoints from a stochastic training trajectory. DSI targets a checkpoint-mixture distribution that averages checkpoint-specific tail fluctuations rather than relying on a single brittle endpoint. We formalize this mechanism through a finite-budget bias-variance theory. Empirically, across multivariate synthetic processes and high-frequency trading data, DSI substantially reduces tail-estimation error compared to single-checkpoint baselines under fixed simulation budgets, outperforming standard diffusion and state-of-the-art tail-aware baselines without modifying the generative objective.

In authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay detection modules. This work shows that by playing our crafted adversarial perturbation as a separate source when the adversary is speaking, the the practical speaker verification system will misjudge the adversary as a target speaker. A two-step algorithm is proposed to optimize the universal adversarial perturbation to be text-independent and has little effect on authentication text recognition. We also estimated room impulse response (RIR) in the algorithm which allowed the perturbation to be effective after being played over the air. In the physical experiment, we achieved targeted attacks with a success rate of 100%, while the word error rate (WER) on speech recognition only increased by 3.55%. And recorded audio could pass replay detection for the live person speaking.
Please feel free to contact me for any questions.