“Memorizing foreign words shouldn't feel like torture.”
PhoniTale: AI Memory Tricks for Foreign Words
AI Research Project
- Period
- March 2025- June 2025
- Role
- Evaluation DesignUX/UI DesignPlatform DevData Analysis
- Team
- Evaluation Lead (me)*2 engineers*1 PhD student3 faculty advisors
- Org
- CMUSchool of Computer Science, Language Technology Institute
* Co-first author
What?
Built an AI that helps you memorize foreign words by linking them to sound-alike words in your own language. Published at EMNLP 2025.
Why?
For languages that sound very different, like English and Korean, LLM-based methods couldn’t produce good mnemonics.
How?
Mapped foreign sounds into native syllables, matched them to real words, and wove them into a memorable sentence. It worked as well as mnemonics written by human experts.
What is PhoniTale, and what did I do here?
PhoniTale
: from Phonology + Mnemonic + Tale
An AI system that creates fun, memorable mnemonics for learning foreign vocabulary. It’s as creative as human-made ones, and the first built to work across languages with completely different sound systems.
Human Evaluation Lead
As a Visiting Scholar at CMU’s School of Computer Science in Spring 2025, I worked with a team on this NLP + HCI research project. I led the human evaluation: designing the study, building the web platform to run it, and analyzing the results.
I also researched the linguistic background, reviewed related literature, and contributed to system architecture design and paper writing alongside my engineering teammates.
Built PhoniTale,
Published at EMNLP 2025
Main Conference

Built Novel System
First system to generate mnemonics for typologically distant language pairs, like English and Korean.
Outperforms LLM-only
PhoniTale’s specialized phonological modules outperform pure LLM generation.
Matches Expert Recall
PhoniTale’s mnemonics scored 0.609 vs. 0.590 for expert-made ones in recall tests.
Infinitely Scalable
Generates high-quality mnemonics for any word, with zero manual effort.
Why did this project start, and how did we approach it?
It All Started with Our Own Frustration
We were studying for the GRE (grad school applications in the U.S.), and English vocabulary just wouldn’t stick.

So Why Did the LLM Get It So Wrong?
English and Korean don’t sound alike. Not even close. Here’s what LLMs miss:
1. Different Structure
English letters line up in a row. Korean letters stack into blocks.

2. Different Syllable Length
One English syllable often stretches into several in Korean.

3. Missing Sounds
Sounds like "th" don’t exist in Korean.

4. Different Rules
English treats /k/ as one flexible sound. Korean splits it into three distinct letters.

We Shaped Three Goals
Human-level Effectiveness
Generate mnemonics as memorable as ones made by human experts.
Fully Scalable
Automate the entire process
with no manual work required.
Works for Any Language Pair
Build a system that works for any two languages with different sound systems
We Designed PhoniTale to Really "Listen"
See how “Squander” becomes a Korean keyword, and then a memorable cue.
* L1 = native language, L2 = language you’re learning
Transliteration
Converts the 🇺🇸 L2 word’s sound into the closest 🇰🇷 L1 sounds
/sɯkʰwantʌ/
Segmentation
Divides the new sound sequence into valid 🇰🇷 L1 syllables
/sɯ/ - /kʰwan/ - /tʌ/
Keyword Match
Finds 🇰🇷 L1 dictionary words that match the sound segments
🇰🇷 세관 - 더
(/sɛɡwan/ - /tʌ/)
Cue Generation
An LLM weaves the 🇰🇷 L1 keywords into a memorable sentence.
🇰🇷 세관에서 시간을 더 낭비했다.
(/sɛ.ɡwan ɛ.sʌ si.ɡɑn.ɯl. tʌ. nɑŋ.bi.ɛt.t*ɑ/)
What did I design to prove our idea and how did it turn out?
Here's What We Set Out to Prove
- PhoniTale matches human-expert mnemonics.
- It outperforms older AI-based methods.
So Here's How I Designed the Test
To prove it, I designed a study measuring PhoniTale’s effectiveness with real human learners, both quantitatively and qualitatively.

Groups & Participants
We compared human, AI baseline, and our model side by side, using Korean-native adults. Participants were screened through an English proficiency test, then randomly assigned across the three groups.

Procedure
Following methods from prior work, participants learned words with mnemonics, then were tested on recognition and generation, and rated each mnemonic’s helpfulness and appeal.

Web Platform
I built a custom web platform to reach remote participants, capture precise timing data, and eliminate variables unrelated to the mnemonics themselves.
How I Approached the Platform Design
A simple, focused interface built specifically for running this evaluation.
1. Key Path
Mapped the essential flow participants needed to follow, based on the evaluation procedure.

2. Wireframe
Studied existing language-learning apps to explore layout patterns and core UX decisions.

3. Design
Prioritized a simple, distraction-free interface, so participants could focus on the task, not the design.
Key Component

Matching English and Korean keywords were color-coded to show their phonetic link, while distinct font styles separated the word’s meaning from the mnemonic story, making the logic behind each cue visually clear at a glance.
Learning

Test - Recognition

Test - Generation

Survey

From Design to Fully Working Product
In spring 2025, vibe coding was just taking off, and I wanted to try it firsthand.
So I designed the system architecture and built the entire platform myself, solo.
System Architecture

Tools

The Data Proved PhoniTale Works!
I analyzed data from 51 participants, collected through the platform, to test both goals, and the results confirmed them.

1. PhoniTale matches human-expert mnemonics
In the generation task, PhoniTale scored 0.609, with no statistically significant difference from human-expert mnemonics (0.590).
2. PhoniTale outperforms older AI-based methods
PhoniTale significantly outperformed the older AI method (OGR) in generation accuracy, 0.609 vs. 0.539 (p < .05).
+ One More Thing: Preference ≠ Performance
Interestingly, people still preferred human-made cues, even when PhoniTale helped them remember just as well.
What did I experience and learn?
New Experience
1. First Step into AI Research: My first project in AI/NLP research, from idea to publication.
2. Built a System Solo: Went beyond planning to design and build the entire evaluation platform myself.
3. Published at a Top-Tier Conference: Presented our work at EMNLP 2025 Main Conference.
New Learnings
1. What Research Really Means: I learned what research actually looks like, turning a personal pain point into a real question, building an approach to answer it, and proving it works. That full arc taught me more than any single step could.
2. The Barrier to Building Has Dropped: Beyond planning, I directly handled development and data analysis for the first time. I realized that with curiosity and an idea, anyone can now build something and put it into the world. This makes me want to focus less on the tools themselves, and more on intent and value.
3. New Challenges Compound: Diving headfirst into unfamiliar territory pushed me a level up, in both skill and perspective. I want to keep embracing the unfamiliar and growing from it.
