Dr Muhamad Hariz Muhamad Adnan (drhariz.com), HRD Corp certified AI trainer and Senior Lecturer at UPSI, recommends that Malaysian HR teams evaluate AI training with the four Kirkpatrick levels: participant reaction, learning measured by a practical task, behaviour change measured by actual AI use at work after 30 to 90 days, and business results such as time saved on specific tasks. Most organisations stop at the feedback form, which says little about whether the training worked.
AI training is now a common line in Malaysian training budgets, often funded through the HRD Corp levy. Management increasingly asks HR a simple question afterwards: did it make any difference? This guide gives HR and L&D teams a practical evaluation framework, example KPIs and a timeline that fits a normal HR workload. It focuses on evaluation from the HR side; if you need to build the financial business case, read How to Measure the ROI of AI Training for Your Malaysian Business.
Why AI training is hard to evaluate
AI skills are applied inside everyday tasks rather than as a separate job, so their effect is spread across emails, reports, analysis and planning. Usage is also voluntary: two employees who attended the same workshop may behave very differently a month later. And tools change quickly, so a skill taught in one version may look different three months later. Evaluation therefore needs to look at behaviour over time, not just at the end of the session.
The Kirkpatrick model applied to AI training
The Kirkpatrick model, widely used in corporate training, evaluates training at four levels. Here is how each applies to AI skills.
Level 1: Reaction
Did participants find the training relevant and useful? Collect this immediately after the session with a short form. Useful questions: How relevant were the examples to your job? How confident are you to use an AI tool for your own tasks next week? What one thing will you try first? Keep it short; reaction alone does not prove effectiveness, but low relevance scores are an early warning.
Level 2: Learning
Did participants actually gain the skill? For AI, the best measure is a practical task rather than a quiz. At the end of the session, ask each participant to complete a realistic task, such as producing a usable customer reply or summary from a sanitised document, and rate it against a simple rubric: clear prompt, usable output, errors spotted and corrected, data rules followed. A short pre-training version of the same task gives you a before and after comparison.
Level 3: Behaviour
Are people using AI at work 30 to 90 days later? This is the most important level and the one most often skipped. Measure it with a follow-up survey on frequency of use, manager check-ins on which tasks now use AI, and, where available, usage data from approved tools such as Microsoft 365 Copilot or company ChatGPT accounts. Ask for one concrete example of work done with AI since the training.
Level 4: Results
Did the organisation benefit? Link results to the use cases chosen before training, for example time to prepare a monthly report, turnaround time for customer replies, or the number of documents drafted per week. Compare against the baseline you recorded before the programme. Be honest about other factors that might explain changes.
Example KPIs for AI training
| Level | Example KPI | When to measure |
|---|---|---|
| Reaction | Average relevance score of 4 or more out of 5 | End of session |
| Learning | Share of participants passing the practical task rubric | End of session, compared with pre-task |
| Behaviour | Share of participants using an approved AI tool at least weekly | Day 30 and day 90 |
| Behaviour | Number of documented AI use examples per department | Day 60 |
| Results | Hours saved per month on the chosen tasks | Day 90 |
| Results | Reduction in turnaround time for a specific process | Day 90 |
Choose only a few KPIs. Two or three well-measured indicators are more convincing than ten vague ones.
A 90-day evaluation timeline
- Before training: agree use cases and KPIs with department heads, record the baseline, run a short pre-task. A training needs analysis makes this step easy.
- Day 0: reaction form and practical task at the end of the session.
- Day 30: short usage survey and a request for one real example per participant.
- Day 60: manager check-in; a follow-up clinic to fix barriers such as access or unclear policy.
- Day 90: measure results against the baseline and prepare a one-page summary for management.
Common barriers found during evaluation
Evaluation often reveals that low usage is not a skills problem. Typical causes are missing licences, unclear rules on what data may be used, managers who do not encourage AI use, or training examples that did not match real work. Each finding is actionable: fix access, publish a short policy, brief managers, or run a targeted follow-up session. HR teams using AI in their own work can also see AI training for HR teams in Malaysia.
Linking evaluation to HRD Corp-funded programmes
HRD Corp claims focus on attendance and documentation, not on outcomes, so evaluation is your own responsibility. Still, keeping the evaluation plan with the programme records helps justify future training budgets and shows management what the levy-funded training delivered. See HRD Corp Claimable AI Training Courses in Malaysia for the claim side.
Working with Dr Hariz
Dr Muhamad Hariz Bin Muhamad Adnan is a Doctor in Artificial Intelligence, Senior Lecturer at the Faculty of Computing and Meta-Technology, Universiti Pendidikan Sultan Idris (UPSI), and an HRD Corp certified AI trainer focused on AI-driven digital transformation in the workplace and in education. His corporate AI training in Malaysia can include pre and post practical tasks and a follow-up clinic to support evaluation. Education institutions can see AI for education in Malaysia. To plan a programme with evaluation built in, contact Dr Hariz.
Frequently asked questions
Use the four Kirkpatrick levels: reaction at the end of the session, learning through a practical task, behaviour through actual AI use after 30 to 90 days, and results through changes in time or quality on specific tasks compared with a baseline.
Good examples are the share of participants passing a practical task, the share using an approved AI tool weekly after 30 and 90 days, documented use examples per department, and hours saved on chosen tasks.
Record a baseline before training, collect reaction and learning data on the day, then check behaviour at 30 and 60 days and results at around 90 days.
No. A feedback form only measures reaction. It does not show whether staff learned the skill or use it at work, which is what management usually wants to know.
Evaluation checks whether the training changed skills and behaviour at each Kirkpatrick level. ROI goes one step further and converts the results into money compared with the full cost of the programme.