TU/e

Testing When a Pedagogical Agent Actually Helps Learning

I designed a 2×2 mixed-factorial study with 157 participants to test whether a visible pedagogical agent helps psychomotor and cognitive learning tasks in the same way. The agent improved psychomotor performance and self-efficacy, but did not improve cognitive performance and slightly reduced cognitive self-efficacy. I also validated AI-based gesture scoring against manual ratings (r=.842) and checked the key findings with robust models when statistical assumptions were not clean.

Role: ResearcherHCI ResearchExperimental DesignHuman-AI InteractionLearning ExperienceQuantitative Research
Research question

The question was not whether a visible agent helps. It was when its presence is worth the attention it takes

Pedagogical agents can demonstrate behavior, provide social cues, and make instruction feel more personal. The same visual presence can also become extra stimulus when the learner is already doing abstract mental work.

For my MSc Human-Technology Interaction thesis at TU/e, I tested whether that tradeoff changes with the task. I compared a visible pedagogical agent with the same instruction delivered without the agent across psychomotor and cognitive Tetris-based tasks.

I designed the experiment around an interaction rather than a universal winner. If the agent was useful because learners could model what they saw, its effect should be stronger when the task actually required physical imitation.

A visible agent should earn its place through the task it supports, not simply because an interface can show one.
Study snapshot

A mixed-factorial experiment separated agent presence from task demand

Agent visibility was between participants. Each participant completed both task types in randomized order.

157participantslab study sample
2×2mixed factorial designvisibility × task type
Randomizedagent conditionvisible or absent
4outcomes measuredperformance, self-efficacy, recall, affective beliefs
Result snapshot

The agent helped most when learners had something physical to model

The headline result was an interaction: visibility did not help every task in the same way.

5.59 vs 4.39psychomotor performancevisible vs absent · p < .001
6.31 vs 5.64psychomotor self-efficacyvisible vs absent · p < .001
n.s.recall interactionagent visibility × task type · p = .703
Task design

The same Tetris rules were adapted into two different kinds of work

Keeping one learning domain helped isolate a more useful question: what changes when the learner has to imitate movement rather than reason through a problem?

Psychomotor task

Learn through hand movement

Participants used predefined hand gestures to control Tetris pieces. A visible agent could provide a behavioral model that was directly relevant to the action being learned.

The visual demonstration was part of how the task could be understood and imitated.
Cognitive task

Learn through mathematical problem-solving

Participants solved mathematical problems to perform game actions. The task depended more on abstract reasoning, making an additional visual agent less directly relevant to the work itself.

This let the study test whether visual presence becomes less useful when modeling is not the main need.
Experiment architecture

The design separated a between-person condition from a within-person task comparison

Participants were randomly assigned to one visibility condition, then completed both cognitive and psychomotor tasks in randomized order.

Sample
Participants enter the lab study

157 participants completed the experiment.

Between-subject
Randomize agent visibility

Each participant is assigned to one visibility condition.

Condition
Agent visible

Instruction includes the on-screen pedagogical agent.

Condition
Agent absent

The same instructional content is shown without the visible agent.

Within-subject
Complete both task types

Cognitive and psychomotor tasks are completed in randomized order, with basic and difficult levels.

Outcomes
Measure learning outcomes

Task performance, self-efficacy, recall, and affective beliefs are analyzed across visibility and task type.

Flow connections

  • participants to assign
  • assign to visible, visible
  • assign to absent, absent
  • visible to tasks
  • absent to tasks
  • tasks to outcomes

The diagram is a portfolio view of the experimental structure. It keeps the between-subject visibility condition separate from the within-subject task comparison.

My scope

I owned the study from experimental structure through analysis

The project combined HCI research design with experimental materials, automated measurement, and quantitative analysis.

Designstructured the 2×2 studyvisibility condition, task types, measures, procedure
Materialsadapted the learning tasksTetris-based cognitive and psychomotor conditions
Measurementvalidated AI-assisted scoringchecked automated gesture scores against manual ratings
Analysistested the interactiontraditional and robust statistical analyses
Experiment materials

The study kept the learning domain stable while changing task demand and agent visibility

These original project artifacts show how the factorial design and the Tetris-based learning tasks were represented in the study.

One experiment crossed visibility with task type

The design compared a visible versus absent agent while participants completed psychomotor and cognitive learning tasks.

Study design for pedagogical agent visibility and task type
Original study-design artifact from the thesis project.
Measurement system

The psychomotor condition needed a scalable score without treating AI as ground truth

Psychomotor performance depended on whether a participant performed specific hand movements correctly. The sessions were video recorded and an AI-based gesture-recognition system was used to score those performance assignments.

Automating the score only helped if it behaved closely enough to manual assessment. I therefore treated validation as part of the measurement design rather than assuming the model output was correct because it was automated.

The AI score was a measurement instrument. It still needed evidence that it tracked the human reference closely enough for the study.
AI scoring validation

Automated gesture scores were checked against manual scoring before being trusted

A validation sample compared the AI-generated scores with manual ratings of the same performance data.

N=152validation scoresAI and manual scoring comparison
r=.842Pearson correlationstrong positive relationship
p=.398mean differencemanual and AI means were not significantly different
0.684MAE / RMSEreported error magnitude for the validation set
Task performance

Visibility improved psychomotor performance, not cognitive performance

Mean task-performance scores by task type and agent condition.

EXPERIMENT · Master thesis · task-performance analysis · 2024 · Psychomotor: N=78 visible / 74 absent · Cognitive: N=80 visible / 77 absent
Visible agentNo agent
01.73.55.27CognitivePsychomotorTask typeMean performance score
TAKEAWAY

The important result is the interaction. The visible agent materially helped the psychomotor task, while cognitive performance did not significantly improve.

View data
Task typeVisible agentNo agent
Cognitive6.136.47
Psychomotor5.594.39

Visibility × task type interaction: F(1,150)=10.72, p<.0013. Cognitive visible vs absent was not significant (p=.27); psychomotor visible vs absent was significant (p<.001).

Self-efficacy

The same interaction appeared in how capable learners felt

Mean self-efficacy ratings by task type and visibility condition.

EXPERIMENT · Master thesis · self-efficacy analysis · 2024 · Reported task-level observations: N=160 visible / 154 absent
Visible agentNo agent
01.73.45.16.8CognitivePsychomotorTask typeMean self-efficacy
TAKEAWAY

Visible agents increased psychomotor self-efficacy, but cognitive self-efficacy moved in the opposite direction. Presence was not uniformly beneficial even when the content stayed the same.

View data
Task typeVisible agentNo agent
Cognitive5.665.97
Psychomotor6.315.64

Visibility × task type interaction: F(1,465)=9.73, p<.0022. Psychomotor visible vs absent p<.001; cognitive visible vs absent p=.043, with the visible condition lower.

What else moved

The broader pattern reinforced a task-dependent interpretation

Null results and affective outcomes matter because they stop the conclusion from becoming 'show an agent everywhere.'

n.s.recallno significant visibility × task interaction · p=.703
p<.001psychomotor content likinghigher with a visible agent; cognitive difference was not significant
5.50 vs 4.81psychomotor instructor likingvisible vs absent · p<.001
Statistical robustness

When the assumptions were messy, I checked whether the conclusion survived a different analysis

Several outcome distributions contained outliers or did not meet normality and variance assumptions cleanly. Relying on one conventional ANOVA result would therefore give more confidence than the data justified.

I used the traditional analyses, then checked key conclusions with robust mixed-effects models. The central interaction for task performance and self-efficacy remained: the effect of visibility depended on the type of task.

The research claim came from a pattern that survived a robustness check, not from one convenient p-value.
Design implication

Agent visibility should behave like a task-level design decision, not a default feature

The study supports a more conditional rule for learning interfaces with visible AI or instructional characters.

Where presence earns its space

Behavioral modeling and physical imitation

When learners need to observe and reproduce an action, a visible model can provide task-relevant information. In this study, that was where performance, self-efficacy, and affective responses improved most clearly.

Use presence because it contributes to the task.
Where presence needs restraint

Abstract reasoning without a modeling need

The visible agent did not improve cognitive task performance and cognitive self-efficacy was slightly lower. Extra visual or social presence should therefore justify what it adds before occupying attention.

Do not equate visibility with support.
Evidence boundary

This was a controlled learning study, not proof that one agent pattern works everywhere

The experiment used Tetris-based tasks in a controlled lab setting. Self-efficacy and affective beliefs were self-reported, outliers were retained, and the AI scoring validation applies to this measurement setup rather than every gesture-recognition system.

Cognitive load theory helped explain why an additional visual model might be less useful during abstract reasoning, but the four dependent variables reported in the study were task performance, self-efficacy, recall, and affective beliefs. I therefore treat cognitive load as an interpretive lens here, not as a directly measured outcome.

A stronger next study would test other learning domains, longer-term exposure, and different agent appearances or behaviors, while continuing to validate any automated performance measure against a human reference.