
About
Hi, I am San Kim, M.S. student in Artificial Intelligence at Sungkyunkwan University (HLI Lab, advisor Prof. JinYeong Bak), researching NLP and language models. Current focus: diffusion language models.
✉️ Email
💻 Github
🔗 LinkedIn
📍 Sungkyunkwan University · HLILab
Education
- M.S. in Artificial Intelligence — Sungkyunkwan University · Mar 2026 – Present (exp. Aug 2028)
- HLI Lab · Advisor: Prof. JinYeong Bak
- B.S. in Computer Science & Engineering — Sungkyunkwan University · Mar 2020 – Dec 2025
Publications
- Evidence-Consistent Fraud Detection Beyond Random Splits
- CIKM 2026, Full Research Paper · Accepted (poster) arXiv
- A decoder-based generative framework (ECoG) that jointly predicts a label, evidence spans, and a rationale for Korean phishing detection under scenario-level OOD splits. → +3.22 Macro-F1, −4.22 prediction–rationale inconsistency, +8.38 evidence-span overlap (0.5B backbone).
- Zero-Shot Detoxification via Multi-Head Attention Scaling
- KCC 2025 · Accepted (co-first author, oral) · DBpia
- A training-free method scaling attention heads by a Positive–Negative Difference (PND) metric to suppress toxic generation. → toxicity 84.2% → 60.7% with lower perplexity (GPT-2 Large).
Research & Projects
HLI Lab, SKKU — Research Intern → M.S. Researcher · Aug 2024 – Present
Advisor: Prof. JinYeong Bak · hli.skku.edu
- Diffusion Language Models — current focus · Jun 2026 – Present
- Ongoing research on diffusion-based language models.
- Voice-Phishing Detection → CIKM 2026 (ECoG) · Accepted · Jul 2025 – May 2026
- Reframed Korean SMS / voice phishing detection as scenario-level OOD generalization (hold out Finance / Parcel / Credit / Government scenarios).
- Decoder-based framework generating label + evidence spans + rationale, with auxiliary consistency objectives discarded at inference.
- Curated a two-modality (SMS + ASR voice) Korean dataset; benchmarked HyperCLOVA X SEED & Qwen3 SLMs.
- Korean Bar Exam RAG · Mar 2025 – Jun 2025
- Evaluated open- and closed-source LLMs under zero-shot, few-shot, and RAG (BM25 / Pyserini over ~150K precedents + ~220K statutes); verified & corrected benchmark labels.
- Diagnosed RAG failure modes (weak retrieval, no re-ranking); proposed GraphRAG, neural re-ranking, and Korean legal-domain pretraining.
- Diary-based Suicide-Risk Prediction (PHQ-9) · Nov 2024 – Feb 2025
- Fine-tuned XLM-RoBERTa via transfer learning; handled a 144-sample constraint with 5-fold CV and ensembling.
Zero-Shot Toxicity Control (Multi-Head Attention Scaling) · Nov 2024 – May 2025
Independent project — KCC 2025 (oral, co-first author)
- Identified toxicity-driving heads via a PND metric; scaled them (α, β) at inference — no retraining, on GPT-2 Large.
- Beat prompt-based baselines on RealToxicPrompts with lower perplexity.
ChartWiz — Visual Instruction Tuning · Jun 2024 – Aug 2024