Evaluating MedEdMentor AI for Theory Recommendation
Can AI help researchers select theories?
You can read this full paper on medRxiv. Highlights are below.
Introduction
On November 9, we created MedEdMENTOR AI, the first AI mentor for medical education research. We have a lot more planned for our AI, but before moving forward we first wanted to evaluate its current performance.
Working with the incomparable Adam Rodman, we completed a unique evaluation to demonstrate the quality of MedEdMENTOR AI. Read on for some of our results and our theory recommendation prompt!
(Again, you can read the full paper on medRxiv)
The task
Utilizing a postpositivist paradigm, we wanted to measure MedEdMENTOR AI's performance on one critical task: recommending theories that medical educators could use to frame a research phenomenon.
This task is important because using proper theoretical framing of education projects facilitates investigation of key aspects of problems and clearer communication of findings across diverse learning contexts.
Evaluation
Now, the interesting thing here is that there's no "right" answer.
Instead, the applicability of a theory is a subjective determination that depends on a researcher’s preferred paradigm, identity, and life experiences.
So then (keeping our postpositivist hats on), how can we measure the quality of MedEdMENTOR AI's theory recommendations?
Creating blinded phenomenon-theory pairs
We utilized a GPT-4 powered literature extraction method to create blinded theory-phenomenon pairs from actual research studies completed in medical education (see the full paper for more details).
We obtained 53 pairs from the last 6 months of medical education literature (June-November 2023). The pairs looked like:
- Research phenomenon - The conceptualization and application of a framework that examines the interplay of various social identities and their positioning within power structures in the context of medical education.
- Actual theory used - Intersectionality
Evaluating the large language models (LLMs)
So now you can see the plan. We could now give a blinded phenomenon to MedEdMENTOR AI and ask it to recommend theories to study it. Then we could compare the recommendations against what was actually used in the research study.
Using the below prompt (based on chain of density), we then asked LLMs to recommend theories to study a given research phenomenon.
You are an expert in medical education research. I will provide you with a phenomenon. Please think deeply about this specific phenomenon, and give me nuanced education theories that may help to clarify the underlying mechanisms pertaining to the phenomenon. When I say nuanced, I mean to think of theories that apply to this phenomenon but not to other medical education research phenomena.
(1) Return 5 nuanced and specific education theories.
(2) Return 5 theories that are more nuanced and specific.
(3) Return 5 theories that are even more nuanced and specific.
Do not explain your answers.
(4) Looking across your entire list, select the most directly applicable 5 theories.
We asked MedEdMENTOR AI and an unmodified "Vanilla" GPT-4 to recommend theories given JUST the phenomenon. (In our paper we also asked a modified MedEdMENTOR AI)
Then we manually reviewed whether the theory recommendations matched what was actually used in the literature.
To summarize briefly: we used research phenomenon from actual studies and asked MedEdMENTOR AI to recommend theories to use. Then we compared those recommendations against the actual theory used to see how many matches there were.
Results
Here are the results:
Performance
- “Vanilla” GPT-4 — 26 of 53 (49%) answers contained a match to the actual theory used in publication.
- MedEdMENTOR AI — 29 of 53 (55%) answers contained a match to the actual theory used in publication.
Example LLM answers
- Research phenomenon - The conceptualization and application of a framework that examines the interplay of various social identities and their positioning within power structures in the context of medical education.
- Actual theory used - Intersectionality
- MedEdMENTOR AI recommendation - Intersectionality Theory, Critical Race Theory, Positionality Theory, Feminist Pedagogy, Transformative Learning Theory
- Research phenomenon - The negotiation of tasks and competencies among healthcare students working together in an interprofessional team during clinical placements.
- Actual theory used - Communities of Practice
- MedEdMENTOR AI recommendation - Interprofessional Education Framework, Communities of Practice, Zone of Proximal Development, Activity Theory, Team-Based Learning Theory
Discussion
These results show MedEdMENTOR AI's promise in the task of selecting an appropriate theory for a medical education phenomenon.
A word about performance:
- As we went through the study, we gained the most improvement from our theory recommendation prompt. We went from only superficial-level recommendations (e.g. repeated recommendations for Andragogy and Cognitive Load Theory) to highly nuanced and specific theories with excellent matching for what humans actually chose.
- Then we made a modest but measurable gain when we provided MedEdMENTOR AI with background documentation on this specific task.
Conclusion
Our experience with MedEdMENTOR AI suggests that LLMs are powerful tools to augment the theoretical constructs of human researchers.
You can see in the prompt above one key approach: we believe that it's key for MedEdMENTOR AI provide a list of theories to researchers, so that each one can be deeply examined for a potential fit with their world view. There is no one right answer for a phenomenon.
Using the theory selected by the authors as the “gold standard” for evaluation of the LLM outputs is inherently a limitation of this study. Simply because the LLM did not suggest the theory that was actually chosen by the authors does not mean that the LLM is “wrong,” because there isn’t necessarily a “right.”
Nonetheless, achieving a 55% match rate with the actual human-chosen theories demonstrates the power of MedEdMENTOR AI.
Read on with the full paper on medRxiv where we show how a modified MedEdMENTOR AI is leading us towards further task-specific training.
We have a LOT more planned — stay tuned!
Take home points
- MedEdMENTOR AI's recommendations appear to be high quality — When provided with phenomena from research literature, MedEdMENTOR AI's recommendations have a 55% rate of matching the human-chosen theory in the literature study.
- Theory recommendations should be menus — Since there is no "right" theory to study a phenomenon, theory recommendation approaches should provide a menu of options for researchers to deeply evaluate each one