International Conference on AI, Data Science, Cybersecurity, Cloud Architectures, and Software Engineering

Theme: Theme details will be published soon.

22-28, April 2026 Holiday Inn Frankfurt Airport – Neu-Isenburg, Frankfurt, Germany
Back to conference
Pengfei Lyu
Featured Speaker

Pengfei Lyu

Session Speaker

United States

Biography

Biography will be updated soon.

Abstract Title

Pengfei Lyu is a postdoctoral researcher in the Department of Biostatistics and Bioinformatics at Duke University. He received his Ph.D. in Statistics from Florida State University in 2024. His research focuses on imbalanced classification, synthetic data generation, bias correction, and high-dimensional dependent data analysis, with applications spanning medical imaging and complex data settings. He is particularly interested in replicability analysis and the development of rigorous statistical frameworks for assessing reproducibility in large-scale studies. His work also explores diffusion models and their integration with deep learning for synthetic data generation and representation learning. In parallel, he develops machine learning methods for echocardiographic imaging, combining generative modeling with classification to improve disease detection. His broader interests include causal inference and reproducible statistical methodologies for real-world data analysis. Reference: Bias-Corrected Data Synthesis for Imbalanced Learning Imbalanced data, where the positive samples represent only a small proportion compared to the negative samples, makes it challenging for classification problems to balance the false positive and false negative rates. A common approach to addressing the challenge involves generating synthetic data for the minority group and then training classification models with both observed and synthetic data. However, since the synthetic data depends on the observed data and fails to replicate the original data distribution accurately, prediction accuracy is reduced when the synthetic data is na\"{i}vely treated as the true data. In this paper, we address the bias introduced by synthetic data and provide consistent estimators for this bias by borrowing information from the majority group. We propose a bias correction procedure to mitigate the adverse effects of synthetic data, enhancing prediction accuracy while avoiding overfitting. This procedure is extended to broader scenarios with imbalanced data, such as imbalanced multi-task learning and causal inference. Theoretical properties, including bounds on bias estimation errors and improvements in prediction accuracy, are provided. Simulation results and data analysis on handwritten digit datasets demonstrate the effectiveness of our method.