Assessing speaking has always been one of the more demanding aspects of language teaching, not because it is unclear what we are looking for, but because what we are looking at is constantly shifting. Unlike writing, where a text remains on the page and can be revisited, spoken language exists only in the moment. It is shaped by pressure, interaction, confidence, and sometimes even by chance. This makes the act of assessment feel less stable, even when clear criteria are in place. In theory, speaking assessment is structured. Teachers are trained to look at fluency, accuracy, range, pronunciation, and interaction. These categories provide a framework, something to rely on when judging performance. Yet in practice, the boundaries between them are rarely clear. A student may speak fluently but with limited range, or use complex language but hesitate frequently. Another may communicate effectively with simple structures, raising the question of whether effectiveness should outweigh technical precision.
What complicates matters further is the role of context. A student speaking confidently in a familiar classroom may struggle in a formal exam setting, even if their language ability has not changed. Anxiety, time pressure, and the presence of an examiner can influence performance in ways that are difficult to separate from actual competence. As a result, what we assess is often not just language, but how well a learner performs under specific conditions. There is also the issue of interaction, which introduces an additional layer of unpredictability. Speaking is rarely a solo act. It depends on the dynamics between speakers, the ability to respond, to listen, to adapt. In pair or group tasks, one student’s performance can affect another’s, either by supporting it or limiting it. This raises questions about fairness, especially when assessment is meant to reflect individual ability.
At the same time, there is a tendency to treat speaking as something that can be measured with precision, as if it were a fixed skill. In reality, it is more fluid. A student may perform well one day and less effectively the next, not because their level has changed, but because the conditions have. This variability does not make assessment impossible, but it does require a level of awareness that goes beyond simply applying criteria. Another challenge lies in the balance between form and meaning. Traditional approaches often place emphasis on accuracy and range, which are easier to identify and score. However, communication is not only about correctness. A student who can express ideas clearly, even with some errors, may be more effective than one who produces technically accurate but limited responses. Deciding how to weigh these elements is not always straightforward. Teachers also bring their own perceptions into the process, often without realising it. Expectations, previous experience with a student, and even subtle biases can influence judgement. This does not mean that assessment is unreliable, but it does suggest that complete objectivity is difficult to achieve. Awareness of this influence is part of maintaining fairness.
In recent years, the introduction of more standardised criteria, particularly through frameworks like the Common European Framework of Reference for Languages, has helped bring greater consistency to speaking assessment. These descriptors offer a shared language for evaluating performance and reduce some of the subjectivity involved. Even so, applying them in real time, while listening and responding, remains a demanding task. Technology has also begun to play a role, offering tools that can record, analyse, and even score spoken language. While these tools can support consistency, they also raise questions about what aspects of speaking can truly be captured. Pronunciation and fluency may be easier to measure, but interaction, tone, and communicative intent are more difficult to quantify.
All of this points to a broader issue. When we assess speaking, we are not simply measuring language ability. We are observing how language is used in a particular moment, under specific conditions, shaped by a range of factors that extend beyond grammar and vocabulary. This makes speaking assessment less about fixed scores and more about informed judgement. For teachers, this means accepting a certain level of complexity. Clear criteria are essential, but they are not enough on their own. Assessment requires attention to context, awareness of variability, and a willingness to look beyond surface features. It also requires recognising that speaking, by its nature, resists complete control. In the end, the goal is not to eliminate uncertainty, but to manage it. Speaking assessment will never be as stable as written assessment, and perhaps it should not be. Language, when spoken, is alive. It shifts, adapts, and responds. Any attempt to measure it must take that into account.
The question, then, is not whether we can assess speaking perfectly. It is whether we understand what we are actually assessing when a student begins to speak.