Reports

Exploring difficulty in Speaking tasks: an intra-task perspective

Official IELTS research report on IELTS Test - Speaking, published in 2006.

Abstract

The oral presentation task has become an established format in high stakes oral testing as examining boards have come to routinely employ them in spoken language tests. This study explores how the difficulty of the Part 2 task (Individual Long Turn) in the IELTS Speaking Test can be manipulated using a framework based on the work of Skehan (1998), while working within the socio-cognitive perspective of test validation. The identification of a set of four equivalent tasks was undertaken in three phases. One of these tasks was left unaltered; the other three were manipulated along three variables: planning time, response time and scaffolded support. In the final phase of the study, 74 language students, at a range of ability levels, performed all four versions of the tasks and completed a brief cognitive processing questionnaire after each performance. The resulting audio files were then rated by two IELTS trained examiners working independently of each other using the current IELTS Speaking criteria. The questionnaire data were analysed in order to establish any differences in cognitive processing when performing the different task versions.

Results from the score data suggest that while the original un-manipulated version tends to result in the highest scores, there are significant differences to be found in the responses of three ability groups to the four tasks, indicating that task difficulty may well be affected differently for test candidates of different ability. These differences were reflected in the findings from the questionnaire analysis. The implications of these findings for teachers, test developers, test validators and researchers are discussed.

1 INTRODUCTION

In recent years, a number of studies have looked at variability in performance on spoken tasks from the perspective of language testing. Empirical evidence has been found to suggest significant effects resulting from test-taker-related variables (Berry 1994, 2004; Kunnan 1995; Purpura 1998), interlocutor-related variables (O’Sullivan 1995, 2000a, 2000b; Porter 1991; Porter & Shen 1991) and rater- and examiner-related variables (Brown 1995, 1998; Brown & Lumley 1997; Chalhoub-Deville 1995; Halleck 1996; Hasselgren 1997; Lazaraton 1996a, 1996b; Lumley 1998; Lumley & O’Sullivan 2000, 2001; Ross 1992; Ross & Berwick 1992; Thompson 1995; Upshur & Turner 1999; Young & Milanovic 1992).

Skehan and Foster (1997) have suggested that foreign language performance is affected by task processing conditions (see also Ortega 1999; Shohamy 1983; Skehan 1998). They have attempted to manipulate processing conditions in order to modify or predict difficulty. In line with this, Skehan (1998) and Norris et al (1998) have made serious attempts to identify task qualities which impinge upon task difficulty in spoken language. They proposed that difficulty is a function of code complexity, cognitive complexity, and communicative demand. A number of empirical findings have revealed that task difficulty has an effect on performance, as measured in the three areas of accuracy, fluency, and complexity (Skehan 1998; Mehnert 1998; Wigglesworth 1997; Skehan & Foster 1997, 1999; Ortega 1999; O’Sullivan, Weir & ffrench 2001).

2 THE ORAL PRESENTATION

‘Oral presentation’ is advocated as a valuable elicitation task for assessing speaking ability by a number of prominent authorities in the field (Clark & Swinton 1979; Bygate 1987; Underhill 1987; Weir 1993, 2005 Hughes 1989, 2003; Butler et al, 2000; Fulcher 2003; Luoma 2004). Its practical advantages are obvious, not least that it can be delivered in a variety of modes. The telling advantage of this method is one speaker produces a long turn alone, without interacting with other speakers. As such, it does not suffer from the ‘contaminating’ effect of the co-construction of discourse in interactive tasks where one participant’s performance will affect the other’s, so is also more suitable for the investigation of intra-task variation, the subject of this study (Iwashita 1997; Luoma 2004; McNamara 1996; Ross & Berwick 1992; Weir 1993, 2005).

Over the past three decades, oral presentation tasks (also known as ‘individual long turn’ or ‘monologic’ tasks) have become an established format in high stakes oral testing as examining boards have come to routinely employ them in spoken language tests. The Test of Spoken English (TSE) from Educational Testing Service (ETS) in the USA, the International English Language Testing System (IELTS), the Cambridge ESOL Main Suite examinations, and the College English Test in China (the world’s biggest EFL examination) all include an ‘oral presentation’ task in their tests of speaking. In ETS’s TOEFL Academic Speaking Test (TAST) only monologues are used. In the context of the New Generation TOEFL speaking component, Butler et al (2000) advocate testing ‘extended discourse’, arguing that this is most relevant to the academic use of language at the university level. Earlier, Clark and Swinton (1979) found that the ‘picture sequence’ task was one of the most effective techniques in experimental tests which investigated suitable techniques for a speaking component for TOEFL.

Given its importance, it is surprising that over the last 20 years no research articles dedicated to oral presentation speaking tasks per se can be found in the most prominent journal in the field, Language Testing. Similarly, there has been little published research on the long turn elsewhere even in the non-language testing literature (see Abdul Raof 2002). Certainly, very little empirical investigation has been conducted to find out what contributes to the degree of task difficulty within oral

presentation tasks in a speaking test even though such tasks play an important function in high stakes tests around the world.

3 TASK DIFFICULTY

In recent years, a number of studies have looked at variability in spoken performance from the perspective of task difficulty in language testing. Empirical evidence has been found to suggest significant effects resulting from how interlocutor-related variables impact on difficulty in interaction-based tasks (Porter 1991; Porter & Shen 1991; O’Sullivan 2000a, 2000b, 2002; Berry 1997, 2004; Buckingham 1997; Iwashita 1997).

In terms of the study of test task related variables, a number of studies concerning inter-task comparison have been undertaken. These have adopted both quantitative perspectives (Chalhoub-Deville 1995; Fulcher 1996; Henning 1983; Lumley & O’Sullivan 2000, 2001; O’Loughlin 1995; Norris et al 1998; Robinson 1995; Shohamy 1983; Shohamy, Reves & Bejarano 1986; Skehan 1996; Stansfield & Kenyon 1992; Upshur and Turner 1999; Wigglesworth & O’Loughlin 1993) and qualitative perspectives (Bygate 1999; Kormos 1999; O’Sullivan, Weir & Saville 2002; Shohamy 1994; Young 1995). These studies were conducted to investigate the impact on scores awarded for speakers’ performances across the different tasks. O’Sullivan and Weir (2002) report that on the whole, the results of these investigations are mixed, perhaps in part due to the crude nature of such investigations where many variables are uncontrolled, and tasks and test populations tend to vary with each study.

There is less research available on intra-task comparison, where internal aspects of one task are systematically manipulated. This is perhaps surprising as this type of study enables the researcher to more closely control and manipulate the variables involved. Skehan and Foster (1997) suggest that foreign language performance is affected by task processing conditions. They propose that difficulty is a function of code complexity, cognitive complexity, and communicative stress. This view is largely supported by the literature (see, for example, Foster & Skehan 1996, 1999; Mehnert 1998; Ortega 1999; Skehan 1996, 1998; Skehan and Foster 2001; Wigglesworth 1997; Brown & Yule 1983; Crookes 1989). The most likely sources of intra-task variability appear to lie in the three broad areas outlined by Skehan and Foster (1997) mentioned above and appear to be most clearly observed when the following specific performance conditions are manipulated:

  1. Planning time
  2. Planning condition
  3. Audience
  4. Type and amount of input
  5. Response time
  6. Topic familiarity

Empirical findings have revealed that intra-task variation in terms of these conditions has an effect on performance as measured in the four areas of accuracy, fluency, complexity and lexical range (Ellis 1987; Crookes 1989; Williams 1992; Skehan 1996; Mehnert 1998; Wigglesworth 1997; Foster & Skehan 1996; Skehan & Foster 1997, 1999; Ortega 1999; O’Sullivan, Weir & ffrench 2001).

Weir (2005) argues that it is critical that examination boards are able to furnish validity evidence on their tests and that this should include research-based evidence on intra-task variation, ie how the conditions under which a single task is performed affect candidate performance. Research into intra-task variation is critical for high stakes tests because if we are able to manipulate the difficulty level of tasks we can create parallel forms of tasks at the same level and offer a principled way of establishing versions of tasks across the ability range (elementary to advanced). This is clearly of relevance to examination bodies that offer a suite of examinations as is the case with Cambridge ESOL.

4 THE STUDY

This study is primarily designed to explore how the difficulty of the IELTS Speaking paper Part 2 task (Individual Long Turn) can be deliberately manipulated using a framework based on the work of Skehan (1998), while working within the socio-cognitive perspective of test validation suggested by O’Sullivan (2000a) and discussed in detail by Weir (2005).

In this research project, the conditions under which tasks are performed are treated as independent variables. We have omitted the variables type and amount of input and topic familiarity from our study as it was decided that it was necessary to limit the scope of the study. These were felt to be adequately controlled for in the task selection process (described in detail below) in which an analysis of the language and topic of each task was undertaken (by considering student responses from the pilot study questionnaire and from the responses of an ‘expert’ panel who applied the difficulty checklist to all tasks). The variable audience was also controlled for by identifying the same audience for each task variant. The remaining variables are operationalised for the purpose of this study in the following way:

VariableUnalteredAltered
Planning Time1 minuteNo planning time
Planning ConditionGuided (3 scaffolding points)No scaffolding
Response Time2 minutes1 minute

Table 1: Task manipulation

The first of the three manipulations is in response to the findings of researchers such as Skehan and Foster (1997, 1999, 2001), Wigglesworth (1997) and Mehnert (1998) who suggest that there is a significant difference in performance where as little as one minute of planning is allowed. Since the findings have shown that this improvement is manifested in increased accuracy, we expect that the scores awarded by raters for this criterion will be most significantly affected. The second area of manipulation is related to the suggestion (by Foster & Skehan, among others) that the nature of the planning can contribute to its effect. For that reason, students will be given an opportunity to engage in guided planning (by using the scaffolded points) or unguided planning (where these points are removed). Finally, the notion of response time is addressed. Anecdotal evidence from examiners and researchers who have listened to recordings of timed responses suggest that test-takers (particularly at a low level of proficiency) tend to run out of things to say and either struggle to add to their performance, engage in repetition of points already made, or simply dry up. Any of these situations can lead to a lowering of the scores candidates are awarded by examiners. Since the original version of this task asks test-takers to respond for 1 to 2 minutes, it was felt to be important to investigate what the consequences of allowing this wide variation in performance time might be.

The hypotheses are formulated as follows:

  1. Planning time will impact on task performance in terms of the test scores achieved by candidates.
  2. Planning condition will impact on task performance in terms of the test scores achieved by candidates.
  3. Response time will impact on task performance in terms of the test scores achieved by candidates.
  4. Differences in performance in respect of the variables in hypotheses 1 to 3 will vary according to the level of proficiency of test-takers.
  5. The manipulations to each task, as represented in hypotheses 1-3, will result in significant changes in the internal processing of the participants (i.e. the theory-based validity of the task will be affected by manipulating elements of the task setting or demands).

4.1 Aims of the study

To establish any differences in candidate linguistic behaviour, as reflected in test scores, arising from language elicitation tasks that have been manipulated along a number of socio-cognitive dimensions

Since all students complete a theory-based validity questionnaire on completion of each of the four tasks they perform (see Appendix 7), analysis of these responses will allow us to make statements regarding the second of our research questions:

To establish any differences in candidate behaviour (cognitive processing) arising from language elicitation tasks that have been manipulated along a number of socio-cognitive dimensions

4.2 Methodology

As mentioned above, this study employs a mixture of quantitative and qualitative methods as appropriate. The study is divided into a number of phases, described below.

Phase 1: In this phase, a number of retired IELTS oral presentation tasks were analysed by the researchers using a checklist based on Skehan (1996). This analysis led to the selection of a series of nine tasks from which it was hoped to identify at least four that were truly equivalent (see Appendix 1 for the checklist). Readability statistics were generated for each of the tasks (see Appendix 2) in order to ascertain that each task was similar in terms of level of input. In addition to these analyses, a qualitative perspective on the task topics was undertaken. The nine tasks are contained in Appendix 3.

Phase 2: A series of pilot administrations was conducted involving overseas university students at a UK institution. These students were on or above the language threshold level for entry into UK university (ie approximately 6.5 on the IELTS overall band scale). The students were asked to perform a number of tasks and to report verbally to one of the researchers on their experience. From these pilot studies it was noted that the topic of two of the tasks (‘visiting a museum or art gallery and ‘entering a contest’) were considered by many students to be outside their experience and as such too difficult to talk about for two minutes. For this reason, the former was changed to a ‘sports event’ and the scaffolding or prompts rewritten, while the latter was dropped from the study. It was decided at this stage that the eight tasks that remained were suitable, and that these should form the basis of the next phase (these are in Appendix 4).

Phase 3: In this phase of the project, a formal trial of the eight selected tasks (A to H) was undertaken.

4.2.1 Quantitative analysis

A group of 54 students was asked to participate in the trial. Each student was asked to complete four tasks, and to fill in a short questionnaire immediately on completing each task. To ensure that an approximately equal number of students responded to each task, the following matrix was devised. This meant that students were given at random a pack marked Version 1 to 8. These packs contained the rubric for each of the tasks in the pack as well as four questionnaires.

Version 1Version 2Version 3Version 4Version 5Version 6Version 7Version 8
AHGFEDCB
BAHGFEDC
CBAHGFED
DCBAHGFE

Table 2: Make-up of task batches for the trial

The above design resulted in the following numbers of students responding to each task.

TaskNumber of Students
A27
B26
C27
D28
E26
F26
G26
H26

Table 3: Number of students responding to each task

The students performed the tasks in a multimedia laboratory, speaking directly to a computer. Each student’s four responses were recorded and saved on the computer as a single file. These files were later edited to remove unwanted elements (such as long breaks following the end of a task performance or unwanted noise that occurred outside of the performance but was inadvertently recorded). The volume of each file was edited to ensure maximum audibility throughout. The performances of each student were then split up into the four constituent tasks and further edited (ie an indicator of student number and task was inserted at the beginning of the task and a bleep inserted to signal to the future rater that the task was now complete). The order of the files was randomised using a random numbers list generated using Microsoft Excel. Finally, eight CDs were created, each of which contained all of the performances for each task.

These eight CDs were then duplicated and a set was given to each of two trained and experienced IELTS raters who rated all tasks over a one-week period. The resulting score data were subjected to multi-faceted Rasch (MFR) analysis using the FACETS program (Linacre 2003) in order to identify a set of at least four tasks where any differences in difficulty could be shown to be statistically insignificant. (For recent examples of this statistical procedure in the language testing literature see Lumley & O’Sullivan 2005, Bonk & Ockey 2004).

The task measurement report from the FACETS output (Table 4) suggests that Task A is potentially significantly easier than the other seven. In addition, the infit mean square statistic (which indicates that all tasks are within the accepted range) suggests that all of the tasks are working in a predictable way.

Fair-M AverageModel MeasureModel S.E.Infit MnSqInfit ZStdOutfit MnSqOutfit ZStdNTask
5.86-.71.111.101.101A
5.74-.27.111.101.112B
5.69-.11.111.001.003C
5.66-.02.11.8-2.8-24D
5.63.08.12.9-1.9-15E
5.51.45.121.211.116F
5.56.29.111.00.907G
5.57.28.111.001.008H

Table 4: Task measurement report (summary of FACETS output)

Follow-up analysis of the scores awarded by the raters indicates that this difference appears to be of statistical significance only in the case of Tasks G and H (see Appendix 5) which appear to be significantly easier than Tasks A and C. The boxplots generated from the SPSS output (Figure 1) suggest that there is a broader spread of scores for Tasks A and C, though in general the mean scores do not appear to be widely spread.

Report figure

Figure 1: Boxplots comparing task means from SPSS output

The results of these analyses suggest that Tasks A, C, G and H should not be considered for inclusion in the main study, though all of the others are acceptable.

4.2.2 Qualitative analysis

In addition to the quantitative analysis described above, we analysed the responses of all students to a short questionnaire (see Appendix 6) about students’ perceptions of the tasks. For this phase of the study, we focused primarily on their responses to the items related to topic familiarity and degree of abstractness of the tasks. The data from these questionnaires (each student completed a questionnaire for each task) were entered into SPSS and analysed for instances of extreme views - as it was thought that we should only accept tasks in which the students felt a degree of comfort that the topic was familiar and that the information given was of a concrete nature. From this analysis, we made a preliminary decision to eliminate two of the eight tasks: Tasks G and H (Table 5). It was decided to monitor Task C, as students perceived it as being somewhat difficult in terms of vocabulary and grammar - though the language of the task (see Appendix 4) does not appear to be significantly different from that of the other tasks.

TopicInformationVocabularyGrammar
TASK12345123451234512345
A98730988101286101110410
B88621961010148400117611
C2135236973112644189820
D9971251263111133101113400
E788206101000158210148310
F4108316911001176011011400
G381131721241145430116810
H731132731033156500115910

KEY: Topic 1 = Familiar 5 = Unfamiliar Information 1 = Very Concrete 5 = Very Abstract Vocabulary & Grammar 1 = Easy 5 = Difficult

Table 5: Qualitative analysis of the tasks (suggesting that G & H be eliminated)

Based on the two types of analyses, the researchers identified four tasks as being equivalent from the qualitative and quantitative perspectives. These were:

Task BTask E
B. Describe a part-time/holiday job that you have done.
You should say:
How you got the job
What the job involved
How long the job lasted
And explain why you think you did the job well or badly.
E. Describe a teacher who has influenced you in your education.
You should say:
Where you met them
What subject they taught
What was special about them
And explain why this person influenced you so much.
Task DTask F
D. Describe an enjoyable event that you experienced when you were at school.
You should say:
What the event was
When it happened
What was good about it
And explain why you particularly remember this event.
F. Describe a film or a TV programme which made a strong impression on you.
You should say:
What kind of film or TV programme it was (eg comedy)
When you saw it
What it was about
And explain why it made such an impression on you.

Figure 2: Four tasks selected for the main study (Phase 5)

In addition to identifying four tasks that can be considered ‘equivalent’ from as broad a number of perspectives as possible, the early phases of the project also saw the development of a series of theory-based validity questionnaires based on ongoing research at the Centre for Research in Testing, Evaluation and Curriculum (CRTEC) at Roehampton University, London (reported by Akmar Zaina Abidin at the Language Testing Forum, Cambridge, 2003). These questionnaires, which are designed to offer insights into the cognitive processing of the participants before and during test task

performance, are based on Weir (2005) and were piloted during Phase 3 (see Appendix 7 for the four versions developed for use in this project).

During this piloting, a number of minor amendments were made to the original drafts based on qualitative feedback from participants - primarily for reasons of clarity and where the language proved to be beyond the level of participating learners.

Phase 4: The above phases meant that we were able to identify a set of four oral presentation tasks for which we could claim equivalence from both qualitative and quantitative perspectives; to the best of our knowledge, this has not been attempted before in either language testing or SLA research.

In this phase, the resulting tasks were manipulated according to the variables identified in Section IV above. Table 6 shows that this manipulation resulted in four versions of each of the four tasks: Task B remained unchanged, Task D had no planning time, Task E had no scaffolding and Task F required a response time of one minute (instead of two minutes).

TaskNo ChangeNo Planning timeNo Scaffolding1 minute response
BxxX
DxxX
ExxX
Fxxx

Table 6: Manipulation of each task

To ensure that there was no order effect, the following matrix was designed (see Table 7). As described above, in this phase of the study, students were asked to perform four tasks, one of which remained unchanged from the original and the others manipulated in the way described in Table 6. In the matrix in Table 7, each version appears on an equal number of occasions and at each level (ie to be performed first, second, etc).

Version 1Version 2Version 3Version 4
BDEF
DBFE
EFBD
FEDB

Table 7: Setup for test versions for the main study

The tasks used in the study can be seen in Figure 3 below.

Figure 3: Manipulation of the tasks in the main study

Figure 3: Manipulation of the tasks in the main study

Phase 5: In the main part of the study, a total of 74 language students at a range of ability levels performed all four versions of the tasks according to the schedule defined by the matrix in Table 7. The resulting audio files were then edited and saved as individual MP3 files. This was done to avoid any halo effect in the rating process as the four tasks performed by any individual were separated so that raters would not be overly affected by performance on an early task when rating the later tasks. Four CDs were created each containing a randomised set of performances for each task (B, D, E and F). These were rated by two IELTS trained examiners working independently of each other using the current rating criteria and scales for the operational IELTS Speaking Test.

5 RESULTS

The scores from these ratings were then analysed using MFR and the resulting data were used for ANOVA and correlational analysis using the programme SPSS, Version 12. The model used in this MFR analysis takes into account the ability of the candidates, the relative harshness of the raters and the difficulty of the tasks to suggest a score called the Fair Average; Fair Average scores have the additional advantage of being true interval in nature.

This will allow us to make statements regarding the first aim of the study:

  • To establish any differences in candidate linguistic behaviour, as reflected in test scores, to language elicitation tasks that have been manipulated along a number of socio-cognitive dimensions

Since all students complete a theory-based validity questionnaire on completion of each of the four tasks they perform (see Appendix 7), analysis of these responses will allow us to make statements regarding the second of our research questions:

  • To establish any differences in candidate behaviour (cognitive processing) to language elicitation tasks that have been manipulated along a number of socio-cognitive dimensions

The existence (or not) of observable systematic differences across the four tasks will be interpreted in light of our third aim:

  • To create a framework for the systematic manipulation of speaking tasks

5.1 Rater agreement

Before analysing the candidate performance data, it is first necessary to explore the area of inter-rater reliability. In this project, a number of measures will be considered, in order to gain a broad picture of the extent to which the two raters behaved in a consistent and predictable way.

First correlation analysis was undertaken to explore the degree to which the two raters placed the candidates in a similar order. The results of this analysis (Table 8) indicate a significant level of correlation for all comparisons (the more meaningful correlations have been highlighted in the table). The overall agreement, based on the raw data is 0.75, certainly acceptable, though not as high as we would expect to find in an operational test event (where it is usual to expect correlations above 0.8). It is possible that the unnatural nature of the rating process, where each rater was given a set of four CDs each one containing the performances of all candidates for a particular task, may have affected rating.

Fluency & coherence 2Lexical resource 2Grammatical range & accuracy 2Pronunciation 2Overall 2
Fluency & coherence 1.700.696.685.629.738
Lexical resource 1.677.662.662.592.694
Grammatical range & accuracy 1.656.631.668.588.679
Pronunciation 1.583.604.651.589.640
Overall 1.720.703.715.643.750

All correlations significant at the 0.01 level (2-tailed).

Table 8: Correlations between the raters

Another estimate of inter-rater agreement is the degree to which they agree on scores around the critical boundary. A widely recognised threshold boundary for IELTS is an overall band score of 6.5 (ie the level demanded by most universities for entrance, computed from scores on the four skills modules); although operational scores for IELTS Speaking are only reported at the whole band level, it was decided to use 6.5 in the following analysis. Table 9 shows the level of agreement/disagreement between the two raters. The shaded areas of the table indicate the areas in which the two raters agreed. This indicates that they agreed for a total of 78% of the candidates and disagreed on the remaining 22%. The table also suggests that Rater 1 is somewhat harsher than Rater 2.

From these two analyses, we can see that the raters were in broad agreement. As both the correlation between the overall scores and the critical boundary agreement indices are acceptable, we can accept that the scores awarded can be used for additional analysis.

Rater 2 PassRater 2 Fail
Rater 1 Pass4845
Rater 1 Fail20183

Table 9: Critical boundary agreement (boundary = 6.5)

5.2 Score data analysis

Following the tests of rater agreement, the first analysis conducted on the task performance score data involved estimating the correlations between the four tasks. Table 10 shows that the correlation were very high and were all significant at the 0.01 level. It is particularly interesting to see that Task B is most highly correlated with Tasks D and F suggesting that the existence of planning time may not significantly affect task performance. Task D was the same as Task B with the single exception that in Task D there was no planning time available to test candidates. The other interesting suggestion here is that the amount of output expected of the candidate does not appear to have had a significant impact on the score achieved. Task F is the same as Task B except that the candidates are expected to talk for two minutes in the former and for just one minute in the latter.

Correlations

Task BTask DTask ETask F
Task B1.900.871.901
Task D.9001.862.858
Task E.871.8621.862
Task F.901.858.8621

All correlations are significant at the 0.01 level (2-tailed).

Table 10: Correlations between the four tasks

To more fully explore the data from the perspective of variation in performance across the four tasks it was decided to classify each candidate into one of three groups; those who are of High ability (setting the critical boundary at 6.5 and including those at and above it); those who could be considered Borderline cases (here the range is from 6.0 to 6.5); and finally those who would have been categorised as Low ability candidates (scoring less than 6.0). All three of these categorisations were based on performance over the four tasks.

VariableCategoryN
Ability LevelPass19
Ability LevelBorderline Fail27
Ability LevelFail28
TaskOriginal74
TaskNo Planning74
TaskNo Support74
TaskReduced Response74

Table 11: Descriptive statistics of the main study data

The descriptive statistics (see Table 11) show that the relative ability level of the population was quite low, with approximately half of the candidates in the ‘fail’ category and only about 20% clearly achieving 6.5 or above. The results of the ANOVA (Table 12) show that there are significant differences between the four task types and the three ability groups (as we would expect since they were selected based on overall scores averages over the four tasks). There does not appear to be any significant interaction between the ability groups and the task type suggesting the stability of these tasks across ability level. However, significant differences emerge in respect of task and ability as separate variables.

SourceType III Sum of SquaresDfMean SquareFSig.
Corrected Model158.490(a)1114.40858.714.000
Intercept9891.75419891.75440309.670.000
Task4.28731.4295.823.001
Ability151.483275.742308.653.000
task * ability2.5706.4281.745.110
Error69.692284.245
Total10066.500296
Corrected Total228.182295

R Squared = .695 (Adjusted R Squared = .683)

Table 12: ANOVA results from the main study

The post hoc (Bonferroni) analysis (Table 13) suggests that there are differences in the responses and that these are significant for comparisons between the original version of the task and the versions which included no planning time and reduced response time. The actual differences in scores achieved for these tasks are approximately one third and one quarter of a band respectively with the original task proving easier in both cases.

ComparisonMean DifferenceSig.95% CI Lower Bound95% CI Upper Bound
OriginalNo Planning.32(*).001.10.54
OriginalNo Support.15.378-.06.37
OriginalReduced Response.26(*).008.05.48
No PlanningNo Support-.17.234-.39.05
No PlanningReduced Response-.061.000-.27.16
No SupportReduced Response.111.000-.10.33

Based on observed means.* The mean difference is significant at the .05 level.

Table 13: Multiple post hoc analysis (Bonferroni)

Having completed the main analyses, a set of charts was then generated. These consisted of a set of clustered boxplots and a line diagram, both of which were based on averaged scores for each task but with ability group also taken into account.

In the first of these charts (Figure 4) we can see that there is relatively little difference in the range of mean scores achieved by each group for the four tasks. While there is a clear difference between the three ability groups in terms of the mean scores achieved by each group for the different tasks, there is also an apparent difference between the pattern of scores on the four tasks between the High ability group (the ‘pass’ group), the Borderline group and the Low ability group (the ‘fail’ group).

Report figure

Figure 4: Boxplots comparing task mean score by ability group

In the final chart (Figure 5 - see following page) we can now see that the pattern of scoring is relatively similar for the Low and Borderline groups but quite different for the High scoring group. Taken with the significant results found in the ANOVA reported above, this suggests that manipulating tasks may result in more complex effects on difficulty than initially thought. The standard version of the task appears to result in optimum performance for all groups; by contrast, the no-planning version appears to result in systematically lower scores across the three ability groups. The lack of support (or scaffolding) appears to have a greater negative impact on test scores achieved by the High and Borderline groups while at the same time having only a very slight (and certainly non-significant) impact on the Low group who may be at a level of language ability where any changes have little impact on performance. Finally, the reduction in response time appears to have had little impact on the performances of the High and Borderline groups, though it clearly has had a different impact on the Low group, with their mean score at its lowest point.

Estimated Marginal Means of tottask Report figure

Figure 5: Line diagram comparing task mean score by ability group

5.3 Questionnaire data analysis (from the perspective of the task)

For reasons of clarity of analysis and presentation, we will present the results from the three parts of the questionnaires separately. In the first part of the questionnaire, all participants were asked to respond to items related to how they dealt with their initial response to each task version. The results are shown in Table 13 below. These results are based on a series of univariate ANOVAs carried out on the data after the questionnaires had been shown to be working as predicted through factor analysis.

The factor analysis of the data was carried out to find evidence that the questionnaires were producing consistent results. Since the three parts of the instrument had been designed to elicit information on specific aspects of the candidates’ behaviour, it was expected that a factor analysis of the responses should result in identifying background factors that matched the planning. The results of the analysis of Part 1 indicated a very clear two-factor solution, with the first four items loading on Factor 1 (which we suggest indicates a more general background knowledge of speaking test response), while the latter four items load a second factor (which appears to be more task-specific knowledge).

FactorItemComponent 1Component 2
Goal settingI read the task very carefully to understand what was required..104.702
Goal settingI thought of HOW to deliver my speech in order to respond well to the topic..114.748
Goal settingI thought of HOW to satisfy the audiences and examiners..273.643
Goal settingI understood the instructions for this speaking test completely..182.657
Generating IdeasI had ENOUGH ideas to speak about this topic..750.236
Generating IdeasI felt it was easy to produce enough ideas for the speech from memory..813.185
Generating IdeasI know A LOT about this type of speech, i.e., I know how to make a speech on this type of topic..823.180
Generating IdeasI know A LOT about other types of speaking test, e.g., interview, discussion..745.126

Extraction Method: Principal Component Analysis. Rotation Method: Varimax with Kaiser Normalisation. A Rotation converged in 3 iterations.

Table 14: Factor analysis of Questionnaire Part 1 (before speaking)

When this is taken into account, the analysis of the responses to individual items should reflect this two-factor solution.

In the first section, which explores candidates’ awareness of how they might go about responding to the task when in the initial stages of reading and considering their response, we can see that there are a number of significant differences between the tasks and the ability groups (though as with all responses to the questionnaire items there is no interaction between the two variables).

ItemAve.Task TypeAbility Group
1. I read the task very carefully to understand what was required.4.2Less likely for No PlanningLess likely for BORDERLINE group
2. I thought of HOW to deliver my speech in order to respond well to the topic.3.7Less likely for No PlanningNo meaningful differences
3. I thought of HOW to satisfy the audiences and examiners.3.3No meaningful differencesNo meaningful differences
4. I understood the instructions for this speaking test completely.4Less likely for No PlanningMore likely for HIGH group
5. I had ENOUGH ideas to speak about this topic.3.1More likely in Original, least for No Planning & No SupportLess likely for LOW group
6. I felt it was easy to produce enough ideas for the speech from memory.3.1More likely in Original, least for No Planning & No SupportLess likely for BORDERLINE group
7. I know A LOT about this type of speech, i.e., I know how to make a speech on this type of topic.2.9No meaningful differencesNo meaningful differences
8. I know A LOT about other types of speaking test, e.g., interview, discussion.3No meaningful differencesNo meaningful differences
  • = no significant difference found • = significant difference found Note: the Likert scale upon which the Average (column 2) is calculated is from 1-5

Table 15: Univariate ANOVA results for Questionnaire Part 1 (before speaking)

The mean response levels (in the Ave. column) indicate that the candidates are likely to read the instructions carefully, and that they tended to have no problem understanding the task. However, they were less likely to consider the audience (Item 3) or to give much thought to the generation of ideas prior to speaking (Items 5 - 8).

It is interesting to note that there is less likelihood that candidates responding to the No Planning version of the tasks will either read the rubric as carefully as for the other versions or think about how to respond in the same was as they might do for the other versions. However, it should be noted that the low mean response to the first item appears to have been heavily influenced by the Borderline group. Review of the data indicates that no errors in data entry could have led to this, and in the absence of post-test interview data, the reason for the very low response cannot easily be explained.

We can also see that the No Planning task appears to have resulted in candidates failing to fully understand the instructions (not surprising in light of the earlier responses which indicated they may not have read them carefully), though this was not a problem for the High ability group.

In the second part of the section, which focused on generating ideas in the pre-planning stage, candidates indicated that the manipulation of the task appears to have had a significant impact on their ability to produce ideas from their background knowledge. Where the task has been altered in terms of planning time or support offered, the candidates report significantly more difficulty in generating ideas - this is most significant for the Low and Borderline groups. For Items 5 and 6 the pattern of response for the Low group was similar across the four tasks, while both the High and Borderline groups indicated a high likelihood for both the Original task and the Reduced Response version and a low likelihood for the other two versions. Perhaps not surprisingly, in the final pair of items, which link the generating of ideas to what is essentially background knowledge, there are no meaningful differences between the tasks or between the three ability levels.

As with the factor analysis of the first section of the questionnaire, the analysis of the second section suggests that this part of the instrument is also working well (Table 16); note that in this analysis the No Planning task was not included as the candidates were not asked to complete a questionnaire since they had not been given any time for planning. The single exception seems to be Item 7, which loads on two factors, so in the analysis that follows this item has been removed. The six-factor solution reflects the original design.

FactorItemComponent 1Component 2Component 3Component 4Component 5
Time Element1. I thought of MOST of my ideas for the speech BEFORE planning an outline.-.071-.070.222.084.635
Time Element2. During the period allowed for planning, I was conscious of the time..114.171-.067-.059.805
Task Specific Planning3. I followed the 3 short prompts provided in the task when I was planning.-.035.771.167-.061-.107
Task Specific Planning4. The information in the short prompts provided was necessary for me to complete the task.-.118.731-.001.042.156
Task Specific Planning5. I wrote down the points I wanted to make based on the 3 short prompts provided in the task.-.111.602.050.443.118
Linguistic Planning6. I wrote down the words and expressions I needed to fulfil the task.-.110.002.152.730.050
Linguistic Planning7. I wrote down the structures I need to fulfil the task..439.000.162.512.310
Language used when Planning8. I made notes only in ENGLISH.-.758.114-.078.209.022
Language used when Planning9. I took notes only in my own language..785-.056.084.157-.001
Language used when Planning10. I took notes in both ENGLISH and own language..862-.092-.016-.039.044
Organisation11. I planned an outline on paper BEFORE starting to speak.-.057-.082.014-.652.045
Organisation12. I planned an outline in my mind BEFORE starting to speak.-.232-.004-.431.410-.200
Generating & Practicing13. Ideas occurring to me at the beginning tended to be COMPLETE.-.111.265.726.059.016
Generating & Practicing14. I was able to put my ideas or content in good order..040.257.661.243-.066
Generating & Practicing15. I practiced the speech in my mind WHILE I was planning..192-.396.584-.015.246
Generating & Practicing16. After finishing my planning, I practiced what I was going to say in my mind until it was time to start..369-.309.543.000.241

Extraction Method: Principal Component Analysis. Rotation Method: Varimax with Kaiser Normalisation. A Rotation converged in 7 iterations.

Table 16: Factor analysis of Questionnaire Part 2 (planning - excludes Task 2)

The mean responses in Table 17 show an interesting pattern, particularly with the high levels for Items 3, 4 and 5 indicating that candidates tended to rely to a great extent on the bullet-pointed prompts: the high mean for Item 8 (when combined with the low means for Items 9 and 10) indicate that planning tends to be done in the target language (though the Low ability group are more likely to use L1). The low means for Items 11 and 12 suggest that little concern is given to planning an outline before speaking. This appears to contradict Item 5, where candidates say they wrote down the points they wanted to make before speaking. It is possible that they interpreted this as actually making a full plan or script of what to say, though not necessarily on paper. This needs to be clarified before any future administration of the instrument.

In the first part of the section (labeled ‘Time Element’) there is little difference across ability levels, though there appears to be a significant effect for the Reduced response version of the task for the item referring to awareness of time. Since there are only two significant effects for all items related to planning we can deduce that manipulating tasks in the ways adapted here may have a limited impact on the planning phase. These aspects can be summarised as:

  • With reduced response time candidates may feel they are under less pressure and so are less conscious of time when responding

  • Removing support from a task appears to make it more difficult for students to plan their response

  • High level candidates are more likely to rely on the supporting points in a task rubric

  • Low level candidates are more likely to use either their own language only or a combination of the target language and their own language in planning.

  • Low level students are more likely to practise what they are about to say both during and after planning

ItemAveTask TypeAbility Level
1. I thought of MOST of my ideas for the speech BEFORE planning an outline.3.64xNo meaningful differencexNo meaningful differences
2. During the period allowed for planning, I was conscious of the time.3.31Least likely for Reduced ResponsexNo meaningful differences
3. I followed the 3 short prompts provided in the task when I was planning.3.99xNo meaningful differencesxNo meaningful differences
4. The information in the short prompts provided was necessary for me to complete the task.3.78xNo meaningful differencesHIGH group more likely to respond positively
5. I wrote down the points I wanted to make based on the 3 short prompts provided in the task.3.84xNo meaningful differencesxNo meaningful differences
6. I wrote down the words and expressions I needed to fulfil the task.3.35xNo meaningful differencexNo meaningful differences
7. I wrote down the structures I need to fulfil the task.2.4xNo meaningful differenceLOW group more likely to respond positively
8. I took notes only in ENGLISH.4.05xNo meaningful differencexNo meaningful differences
9. I took notes only in my own language.1.9xNo meaningful differenceLOW group more likely to respond positively (but low means)
10. I took notes in both ENGLISH and own language.2.14xNo meaningful differenceLower level more likely to respond positively
11. I planned an outline on paper BEFORE starting to speak.1.25xNo meaningful differencexNo meaningful differences
12. I planned an outline in my mind BEFORE starting to speak.1.38xNo meaningful differencexNo meaningful differences
13. Ideas occurring to me at the beginning tended to be COMPLETE.3.12xNo meaningful differencexNo meaningful differences
14. I was able to put my ideas or content in good order.2.88Less likely for No SupportxNo meaningful differences
15. I practiced the speech in my mind WHILE I was planning.2.89xNo meaningful differenceLOW group more likely to respond positively (but low means)
16. After finishing my planning, I practiced what I was going to say in my mind until it was time to start.2.72xNo meaningful differenceHIGH group less likely to respond positively
  • = no significant difference found • = significant difference found Note: Items 3, 4 and 5 not included in No Support version (as they refer to supporting points)

Table 17: Univariate ANOVA results for Questionnaire Part 2 (during planning)

In the final section of the questionnaire, candidates were asked to respond to items related to what they did as they were speaking. The factor analysis reflected the original design, as so the section was considered to have worked as predicted.

FactorItemComponent 1Component 2Component 3Component 4
Idea Development (ability)1. I felt it was easy to put ideas in good order..819.083.079-.028
Idea Development (ability)2. I was able to express my ideas using appropriate words..705.203.134.015
Idea Development (ability)3. I was able to express my ideas using correct grammar..695.194.133.088
Idea Development (ability)6. I was able to put sentences in logical order..736.226.086.040
Idea Development (ability)7. I was able to CONNECT my ideas smoothly in the whole speech..602.264.073-.136
Idea Development (ability)14. I felt it was easy to complete the task..748.125.158.094
Idea Development (temporal)4. I thought of MOST of my ideas for the speech WHILE I was actually speaking.-.048.205.330.714
Idea Development (temporal)5. Some ideas had to be omitted while I was speaking..103-.132-.326.759
Time Awareness8. I was conscious of the time WHILE I was making this speech..194.009.819-.025
Time Awareness9. I tried NOT to speak more than the required length of time in the instructions..239.278.629.012
Monitoring10. I was listening and checking the correctness of the contents and their order WHILE I was making this speech..251.754.030-.017
Monitoring11. I was listening and checking whether the contents and their order fit the topic WHILE I was making this speech..195.786.049-.020
Monitoring12. I was listening and checking the correctness of sentences WHILE I was making this speech..215.783.090.016
Monitoring13. I was listening and checking whether the words fit the topic WHILE I was making this speech..170.744.221.107

Extraction Method: Principal Component Analysis. Rotation Method: Varimax with Kaiser Normalisation. A Rotation converged in 5 iterations.

Table 18: Factor analysis of Questionnaire Part 3 (during speaking)

The most interesting thing about mean responses in this section is the lack of variation across the items. In the first part, there is very much a ‘no view’ perspective displayed, suggesting that the candidates were not overly challenged by the tasks. In support of the findings for the previous section, there appears to have been a tendency for candidates to plan while speaking (Item 4) and a slight tendency for them to monitor the contents and language of their responses (though the latter seems to have been most likely with the High ability level).

In the first part of the section, which related to ease and ability to develop ideas, the suggestion appears to be that the candidates found the Original version of the task the easiest to respond to (though this was shared with the Reduced Response version for Item 1). Not surprisingly the High level candidates indicated that they found it easy to “express…ideas using good grammar,” while the Borderline candidates seemed to struggle with cohesion and coherence.

Low level candidates were more likely to omit ideas as they were speaking, though this was reported as being less likely with the No Support task version, possibly because the candidates considered the ‘idea’ to be related primarily with the three bullet-pointed supporting points suggested and when these were removed they struggled.

ItemAve.Task TypeAbility Level
1. I felt it was easy to put ideas in good order.2.9Easier for Original and Reduced Response×No meaningful differences
2. I was able to express my ideas using appropriate words.3×No meaningful differences×No meaningful differences
3. I was able to express my ideas using correct grammar.2.8×No meaningful differencesMore likely with HIGH group
6. I was able to put sentences in logical order.3×No meaningful differencesLess likely with BORDERLINE group
7. I was able to CONNECT my ideas smoothly in the whole speech.2.8More likely with Original, especially compared to No PlanningLess likely with BORDERLINE group
14. I felt it was easy to complete the task.2.9×No meaningful differences×No meaningful differences
4. I thought of MOST of my ideas for the speech WHILE I was actually speaking.3.4×No meaningful differences×No meaningful differences
5. Some ideas had to be omitted while I was speaking.3Less likely with No Support versionMost likely for LOW group
8. I was conscious of the time WHILE I was making this speech.3.3×No meaningful differences×No meaningful differences
9. I tried NOT to speak more than the required length of time in the instructions.3.4×No meaningful differences×No meaningful differences
10. I was listening and checking the correctness of the contents and their order WHILE I was making this speech.3.3×No meaningful differences×No meaningful differences
11. I was listening and checking whether the contents and their order fit the topic WHILE I was making this speech.3.3×No meaningful differencesLess likely with Borderline group
12. I was listening and checking the correctness of sentences WHILE I was making this speech.3.3×No meaningful differencesMore likely with HIGH group
13. I was listening and checking whether the words fit the topic WHILE I was making this speech.3.3×No meaningful differencesMore likely with HIGH group
  • = no significant difference found • = significant difference found

Table 19: Univariate ANOVA results for Questionnaire Part 3 (during speaking)

Time did not seem to be particularly important to candidates, and though there was a slight tendency for them to be conscious of time, this does not appear to have varied across ability level or task type attempted. Similarly, though candidates tended to monitor their responses for content, organisation and language, this was not a very strong trend, with the exception of the High ability group who were significantly more likely to monitor their language (but not content or organisation) than the other groups.

6 CONCLUSIONS

In this research project we set out to establish whether the difficulty of a task could be varied by systematic manipulation along a number of dimensions. In doing this we were interested in whether the scores achieved by a group of test candidates would vary along with the cognitive processing associated with performance on the various tasks. This was hoped to provide the basis for a framework which could be used to manipulate tasks in order to systematically alter the difficulty of these tasks.

The project called for a set of four equivalent tasks to be identified so that all participants would respond to an unaltered version as well as three versions in which systematic variations had been made (removal of planning time; removal of support; and reduction of expected response time). In order to identify four equivalent tasks, a complex procedure was designed, in which a set of nine tasks was analysed both quantitatively (based on the performances of a group of 54 participants) and qualitatively (using the responses of these same participants to a series of short questionnaires).

At this stage, a set of four tasks was identified and manipulated as planned. A group of 74 participants then recorded their responses to the tasks which were presented to different people in different orders. At the same time, all respondents then completed questionnaires (one per task, so a total of four per participant) based on Weir’s (2005) socio-cognitive framework for test validation for speaking. The resulting data were then analysed using the two datasets.

Results of the analysis of the score data suggest that there are significant differences to be found in the responses of three ability groups to the four tasks, indicating that task difficulty may well be affected differently for test candidates of different ability. In other words, simply altering a task along a particular dimension may not result in a version that is equally more or less difficult for all test candidates. Instead, there is likely to be a variety of effects as a result of the alteration. For instance, here, mid-level and higher-level participants were not significantly affected by the reduction in response time, while this same alteration to the task resulted in the most serious negative effect for the lower level participants.

The analysis of the questionnaire data further complicates the picture. We can briefly summarise the findings as:

  • The most significant effects of task manipulation on candidates appear to be at the prespeaking phase, particularly where no planning time is offered. However, these effects appear to differ depending on the ability level of the candidate.
  • The effects on planning are far less obvious. The candidates report essentially the same approach to planning regardless of the task. Here, while there are far more significant differences in the ways in which candidates of different ability level approach task planning, there appears to be a clear tendency for them not to outline their response before speaking, so even though they take the time to plan, they seem to do much of their planning ‘on-line’ ie, as they are speaking (though lower level candidates report practising what they plan to say before speaking).
  • When speaking, the candidates seemed to feel that the original version of the task offered them the greatest opportunity to perform at their best, though not surprisingly, this depended on their ability level (lower levels did not find any particular version easier in any way than the others). There was a significant difference in approach to monitoring of own output, with the higher level students more likely to monitor language, though not content or organisation).

6.1 Implications

We believe the study has implications for teachers who prepare students for examinations containing speaking tasks which involve individual long turn responses, for the test developers who design these tasks, for test validators and first and second language acquisition researchers.

6.1.1 Teachers

The differences in approach to task performance highlighted here suggest that teachers might focus more explicitly on pre-speaking strategies such as focusing more clearly on any bulleted prompts and on using the target language for any planning. The lack of impact on approach to planning of task manipulation suggests that students (certainly those involved in this study) have already formed strategies for task performance. However, to improve their understanding of a task, students should be encouraged to read task rubrics more carefully, focus on the language used in the instructions and perhaps ask for assistance where things are not clear.

6.1.2 Test developers

The notion of task equivalence is not as straightforward as it seems. The nine tasks initially used here were presumed by their developers to be equivalent. The methodology used to establish equivalence demonstrated how difficult it can be to create truly equivalent versions of a task. The main study also demonstrates how task difficulty can be affected by decisions to either include or exclude support (eg in the form of bulleted prompts) or by altering the planning time afforded to candidates. This suggests that any substantive changes to these conditions of task performance need to be empirically tested before they are considered in any test revision (or as alternative choices within a test). This is particularly relevant for the planning variable, where the difference in scores achieved was significantly lower for the ‘no planning’ condition than for the original version of the task (which allows one minute of planning time).

The situation regarding amount of response time seems to be less conclusive. Apart from a reduced awareness of time in the planning phase (possibly due to the perception that less speaking time meant there was less to worry about), there appears to have been no difference to the approach taken to task response. However, the scores achieved appear to have been significantly lower for this version than for the original version of the task (in the original version candidates spoke for 2 minutes as opposed to 1 minute in the reduced response version).

The rubric appears to be especially important in this type of task. It is clear that a number of candidates (typically at the lower level) had some difficulty understanding what to do. While this is possibly unavoidable in a test which is designed to be used across a broad range of abilities, it is clearly very important for the test developer to ensure measures are in place to avoid poor reading or listening skills affecting student spoken performance. In ‘live’ tests this is not so difficult (examiners can be trained to deal systematically with comprehension problems), though it is a potentially serious limitation of any computer-delivered test of this sort.

6.1.3 Test validators

In the same way that test developers need to focus on the area of task equivalence, test validators should also consider the area when establishing evidence of the context validity (see Weir 2005) of their tests. Consideration should be given to using the methodology developed here in order to establish true equivalence in test tasks, as well as to investigating how tasks are affected when variations are suggested by stakeholders.

6.1.4 Researchers

SLA researchers have argued since the mid-1980s that performing language elicitation tasks in a learning environment supports learning. While O’Sullivan (2000a: 298) argues that ‘[The] notion of an interlocutor effect on performance does not appear to have been sufficiently addressed in the [SLA] literature’, he also argues that the ’conditions under which tasks are performed should be more rigorously described’ (O’Sullivan, 2000a: 297). While there has been a recognition in the taskbased learning literature that task performance conditions can affect performance (Larson-Freeman & Long, 1991: 30-33), there is little evidence that this awareness has found its way into SLA or Applied Linguistics research.

The evidence presented in this project suggests that researchers need to more clearly understand the implications of decisions they make when designing tasks for use as elicitation devices in their studies. Research studies should contain both more detail of task design and equivalence and an awareness on the side of the researcher of the rationale for task selection and manipulation. In other words, tasks for both testing and research purposes should be specified in an equally systematic and comprehensive fashion using a model of validation such as that of Weir (2005) to ensure that the results obtained are credible in terms of the validity evidence available.

References

  • Abdul Raof, AH, 2002, ‘The production of a performance rating scale: an alternative methodology’, unpublished PhD dissertation, The University of Reading, UK
  • Berry, V, 1994, ‘Personality characteristics and the assessment of spoken language in an academic context’, paper presented at the 16th Language Testing Research Colloquium, Washington, DC
  • Berry, V, 1997, ‘Gender and personality as factors of interlocutor variability in oral performance tests’, paper presented at the 19th Language Testing Research Colloquium, Orlando, Florida
  • Berry, V, 2004, ‘A study of the interaction between individual personality differences and oral test performance test facets’, unpublished PhD dissertation, Kings College, The University of London
  • Bonk, WJ and Ockey, GJ, 2003, ‘A many-facet Rasch analysis of the second language group oral discussion task’, Language Testing, vol 20, no 1, pp 89-110
  • Brown, A, 1995, ‘The effect of rater variables in the development of an occupation specific language performance test’, Language Testing, vol 12, no 1, pp 1-15
  • Brown, A, 1998, ‘Interviewer style and candidate performance in the IELTS oral interview’, paper presented at the 20th Language Testing Research Colloquium, Monterey, CA
  • Brown, A, and Lumley, T, 1997, ‘Interviewer variability in specific-purpose language performance tests’ in Current Developments and Alternatives in Language Assessment, eds A Huhta, V Kohonen, L Kurki-Suonio and S Luoma, University of Jyväskylä and University of Tampere, Jyväskylä, pp137-150
  • Brown, G, and Yule, G, 1983, Teaching the spoken language, Cambridge University Press, Cambridge
  • Buckingham, A, 1997, ‘Oral language testing: do the age, status and gender of the interlocutor make a difference?’, unpublished MA dissertation, University of Reading
  • Butler, FA, Eignor, D, Jones, S, McNamara, T, and Suomi, BK, 2000, TOEFL (2000) Speaking Framework: A Working Paper, TOEFL Monograph Series 20, Educational Testing Service, Princeton, NJ
  • Bygate, M, 1987, Speaking, Oxford University Press, Oxford
  • Bygate, M, 1999, ‘Quality of language and purpose of task: patterns of learners’ language on two oral communication tasks’, Language Teaching Research, vol 3, no 3, pp 185-214
  • Chalhoub-Deville, M, 1995, ‘Deriving oral assessment scales across different tests and rater groups’, Language Testing, vol 12, pp16-33
  • Clark, JLD and Swinton, SS, 1979, ‘An exploration of speaking proficiency measures in the TOEFL context’, TOEFL Research Report, Educational Testing Service, Princeton, NJ
  • Crookes, G, 1989, ‘Planning and interlanguage variation’, Studies in Second Language Acquisition, vol 11, pp 367-383
  • Ellis, R, 1987, ‘Interlanguage variability in narrative discourse: style shifting in the use of the past tense’, Studies in Second Language Acquisition, vol 9, pp 1-20
  • Foster, P and Skehan, P, 1996, ‘The influence of planning and task type on second language performance’, Studies in Second Language Acquisition, vol 18, pp 299-323
  • Foster, P and Skehan, P, 1999, ‘The influence of source of planning and focus of planning on taskbased performance’, Language Teaching Research, vol 3, no 3, pp 215-247
  • Fulcher, G, 1996, ‘Testing tasks: issues in task design and the group oral’, Language Testing, vol 13, no 1, pp 23-51
  • Fulcher, G, 2003, Testing second language speaking, Longman/Pearson, London
  • Halleck, G, 1996, ‘Interrater reliability of the OPI: using academic trainee raters’, Foreign Language Annals, vol 29, no 2, pp 223-238
  • Hasselgren, A, 1997, ‘Oral test subskill scores: what they tell us about raters and pupils’, in Current Developments and Alternatives in Language Assessment, eds A Huhta, V Kohonen, L Kurki-Suonio and S Luoma, University of Jyväskylä and University of Tampere, Jyväskylä, pp 241-256
  • Henning, G, 1983, ‘Oral proficiency testing: comparative validities of interview, imitation, and completion methods’, Language Learning, vol 33, no 3, pp 315-332
  • Hughes, A, 1989, Testing for language teachers, Cambridge University Press, Cambridge
  • Hughes, A, 2003, Testing for language teachers: Second Edition, Cambridge University Press, Cambridge
  • Iwashita, N, 1997, ‘The validity of the paired interview format in oral performance testing’, paper presented at the IQ˙th\mathcal { I } ^ { \dot { Q } ^ { t h } } Language Testing Research Colloquium, Orlando, Florida
  • Kormos, J, 1999, ‘Simulation conversations in oral proficiency assessment: a conversation analysis of role plays and non-scripted interviews in language exams’, Language Testing, vol 16, no 2, pp 163-188
  • Kunnan, AJ, 1995, Test-taker characteristics and test performance: a structural modeling approach, UCLES/Cambridge University Press, Cambridge
  • Larson-Freeman, D, and Long, MH, 1991, An introduction to second language acquisition research, Longman, London
  • Lazaraton, A, 1996a, ‘Interlocutor support in oral proficiency interviews: the case of CASE, Language Testing, vol 13, no 2, pp 151-172
  • Lazaraton, A, 1996b, ‘A qualitative approach to monitoring examiner conduct in the Cambridge Assessment of Spoken English (CASE)’, in Performance testing, cognition and assessment: selected papers from the ISˉth\stackrel { \bullet } { I } \bar { S } ^ { t h } Language Testing Research Colloquium, Cambridge and Arnhem, eds M Milanovic and N Saville, UCLES/Cambridge University Press, Cambridge, pp 18-33
  • Linacre, JM, 2003, FACETS 3.45 computer program, MESA Press, Chicago, IL
  • Lumley, T, 1998, ‘Perceptions of language-trained raters and occupational experts in a test of occupational English language proficiency’, English for Specific Purposes, vol 17, no 4, pp 347-367
  • Lumley, T and O’Sullivan, B, 2000, ‘The effect of speaker and topic variables on task performance in a tape-mediated assessment of speaking’, paper presented at the 2nd2 ^ { n d } Annual Asian Language Assessment Research Forum, The Hong Kong Polytechnic University
  • Lumley, T and O’Sullivan, B, 2001, ‘The effect of test-taker sex, audience and topic on task performance in tape-mediated assessment of speaking’, Melbourne Papers in Language Testing, vol 9, no 1, pp 34-55
  • Lumley, T and O’Sullivan, B, 2005, ‘The effect of test-taker gender, audience and topic on task performance in tape-mediated assessment of speaking’, Language Testing, vol 23, no 4, pp 415-437
  • Luoma, S, 2004, Assessing Speaking, Cambridge University Press, Cambridge
  • McNamara, T, 1997, ‘Interaction’ in second language performance assessment: whose performance?’ Applied Linguistics, vol 18, pp 446-466
  • Mehnert, U, 1998, ‘The effects of different lengths of time for planning on second language performance’, Studies in Second Language Acquisition, vol 20, pp 83-108
  • Norris, J, Brown, JD, Hudson, T and Yoshioka, J, 1998, Designing second language performance assessment, Technical Report #18, University of Hawai’i Press, Hawai’i
  • O’Loughlin, K, 1995, ‘Lexical density in candidate output on direct and semi-direct versions of an oral proficiency test’, Language Testing, vol 12, no 2, pp 217-237
  • O’Sullivan, B, 1995, ‘Oral language testing: does the age of the interlocutor make a difference?’ unpublished MA dissertation, University of Reading
  • O’Sullivan, B, 2000a, ‘Towards a model of performance in oral language testing’, unpublished PhD dissertation, University of Reading
  • O’Sullivan, B, 2000b, ‘Exploring gender and oral proficiency interview performance’, System, vol 28, no 3, pp 373-386
  • O’Sullivan, B, 2002, ‘Learner acquaintanceship and oral proficiency test pair-task performance’, Language Testing, vol 19, no 3, pp 277-295
  • O’Sullivan, B, and Weir, C, 2002, Research issues in testing spoken language, mimeo: internal research report commissioned by Cambridge ESOL
  • O’Sullivan, B, Weir, C and ffrench, A, 2001, ‘Task difficulty in testing spoken language: a sociocognitive perspective’, paper presented at the 23rd Language Testing Research Colloquium, St Louis, Miss
  • O’Sullivan, B, Weir, CJ and Saville, N, 2002, ‘Using observation checklists to validate speaking-test tasks’, Language Testing, vol 19, no 1, pp 33-56
  • Ortega, L, 1999, ‘Planning and focus on form in L2 oral performance’, Studies in Second Language Acquisition, vol 20, pp 109-148
  • Porter, D, 1991, ‘Affective factors in language testing’ in Language Testing in the 1990s, eds JC Alderson and B North, Modern English Publications in association with British Council, Macmillan, London, pp 32-40
  • Porter, D and Shen SH, 1991, ‘Gender, status and style in the interview’, The Dolphin 21, Aarhus University Press, pp 117-128
  • Purpura, J, 1998, ‘Investigating the effects of strategy use and second language test performance with high- and low-ability test-takers: a structural equation modeling approach’, Language Testing, vol 15, no 3, pp 333-379
  • Robinson, P, 1995, ‘Task complexity and second language narrative discourse’, Language Learning, vol 45, no 1, pp 99-140
  • Ross, S, 1992, ‘Accommodative questions in oral proficiency interviews’, Language Testing, vol 9, pp 173-186
  • Ross, S and Berwick, R, 1992, ‘The discourse of accommodation in oral proficiency interviews’, Studies in Second Language Acquisition, vol 14, pp 159-176
  • Shohamy, E, 1983, ‘The stability of oral language proficiency assessment on the oral interview testing procedure’, Language Learning, vol 33, pp 527-540
  • Shohamy, E, 1994, ‘The validity of direct versus semi-direct oral tests’, Language Testing, vol 11, pp 99-123
  • Shohamy, E, Reves, T and Bejarano, Y, 1986, ‘Introducing a new comprehensive test of oral proficiency’, ELT Journal, vol 40, no 3, pp 212-220
  • Skehan, P, 1996, ‘A framework for the implementation of task based instruction’, Applied Linguistics, vol 17, pp 38-62
  • Skehan, P, 1998, A cognitive approach to language learning, Oxford University Press, Oxford
  • Skehan, P and Foster, P, 1997, ‘The influence of planning and post-task activities on accuracy and complexity in task-based learning’, Language Teaching Research, vol 1, no 3, pp 185-211
  • Skehan, P and Foster, P, 1999, ‘The influence of task structure and processing conditions on narrative retellings’, Language Learning, vol 49, no 1, pp 93-120
  • Skehan, P and Foster, P, 2001, ‘Cognition and tasks’ in Cognition and second language instruction, ed P Robinson, Cambridge University Press, Cambridge, pp 183-205
  • Stansfield, CW and Kenyon, DM, 1992, ‘Research on the comparability of the oral proficiency interview and the simulated oral proficiency interview’, System, vol 20, pp 347-364
  • Thompson, I, 1995, ‘A study of interrater reliability of the ACTFL oral proficiency interview in five European Languages: data from ESL, French, German, Russia, and Spanish’, Foreign Language Annals, vol 28, no 3, pp 407-422
  • Underhill, N, 1987, Testing spoken language: a handbook of oral testing techniques, Cambridge University Press, Cambridge
  • Upshur, JA and Turner, C, 1999, ‘Systematic effects in the rating of second-language speaking ability: test method and learner discourse’, Language Testing, vol 1, no 1, pp 82-111
  • Weir, CJ, 1990, Communicative language testing, Prentice Hall International
  • Weir, CJ, 1993, Understanding and developing language tests, Prentice Hall London
  • Weir, CJ, 2005 Language testing and validation: an evidence-based approach, Palgrave, Oxford
  • Wigglesworth, G, 1997, ‘An investigation of planning time and proficiency level on oral test discourse’, Language Testing, vol 14, no 1, pp 85-106
  • Wigglesworth, G, and O’Loughlin, K, 1993, ‘An investigation into the comparability of direct and semi-direct versions of an oral interaction test in English’, Melbourne Papers in Language Testing, vol 2, no 1, pp 56-67
  • Williams, J, 1992, ‘Planning, discourse marking, and the comprehensibility of international teaching assistants’, TESOL Quarterly, vol 26, pp 693-711
  • Young, R, 1995, ‘Conversational styles in language proficiency interviews’, Language Learning, vol 45, no 1, pp 3-42
  • Young, R, and Milanovic, M, 1992, ‘Discourse variation in oral proficiency interviews’, Studies in Second Language Acquisition, vol 14, pp 403-424

Appendix 1: Task Difficulty Checklist (Based On Skehan, 1998)

MODERATOR VARIABLESCONDITIONGLOSS (THE MORE DIFFICULT THE HIGHER THE NUMBER)DIFFICULTY (CIRCLE ONE)
CODE COMPLEXITYRange of linguistic inputVocabulary and structure as appropriate to ALTE levels 1 - 5 (beginner to advanced)1 2 3 4 5 6
CODE COMPLEXITYSources of inputNumber and types of written and spoken input
1 = one single written or spoken source
5 = multiple written and spoken sources
1 2 3 4 5 6
COGNITIVE COMPLEXITYAmount of linguistic input to be processedQuantity of input
1 = sentence level (single question, prompts)
5 = long text (extended instructions and/or texts)
1 2 3 4 5 6
COGNITIVE COMPLEXITYAvailability of inputExtent to which information necessary for task completions is readily available to the candidate
1 = all information provided
5 = student attempts an open ended task [student provides all information]
1 2 3 4 5 6
COGNITIVE COMPLEXITYFamiliarity of information1 = the information given and/or required is likely to be within the candidates’ experience
5 = information given and/or required is likely to be outside the candidates’ experience
1 2 3 4 5 6
COGNITIVE COMPLEXITYOrganisation of information required1 = almost no organisation required
5 = extensive organisation required from a simple answer to a question to a complex response
1 2 3 4 5 6
COGNITIVE COMPLEXITYAs information becomes more abstract1 = concrete
5 = abstract
1 2 3 4 5 6
COMMUNICATIVE DEMANDTime pressure1 = no constraints on time available to complete task (if candidate does not complete the task in the time given he/she is not penalised)
5 = serious constraints on time available to complete task (if candidate does not complete the task in the time given he/she is penalised)
1 2 3 4 5 6
COMMUNICATIVE DEMANDResponse level1 = more than sufficient to plan or formulate a response
5 = no planning time available
1 2 3 4 5 6
COMMUNICATIVE DEMANDScaleNumber of participants in a task, number of relationships involved
1 = one person
5 = five or more people
1 2 3 4 5 6
COMMUNICATIVE DEMANDComplexity of task outcome1 = simple unequivocal outcome
5 = complex unpredictable outcome
1 2 3 4 5 6
COMMUNICATIVE DEMANDReferential complexity1 = reference to objects and activities which are visible
5 = reference to external/displaced (not in the here and now) objects and events
1 2 3 4 5 6
COMMUNICATIVE DEMANDStakes1 = a measure of attainment which is of value only to the candidate
5 = a measure of attainment which has a high external value
1 2 3 4 5 6
COMMUNICATIVE DEMANDDegree of reciprocity required1 = no requirement of the candidate to initiate, continue or terminate interaction
5 = task requires each candidate to participate fully in the interaction
1 2 3 4 5 6
COMMUNICATIVE DEMANDStructured1 = task is highly structured/scaffolded
5 = task is totally unstructured/unscaffolded
1 2 3 4 5 6
COMMUNICATIVE DEMANDOpportunity for control1 = complete autonomy
5 = no opportunity for control
1 2 3 4 5 6

Appendix 2: Readability Statistics For 9 Tasks

Task 1Task 2Task 3Task 4Task 5Task 6Task 7Task 8Task 9
Counts
Words353336433435463138
Characters153142150162169169185146151
Paragraph111111111
Sentences666666666
Average
Sentence/Paragraph6.06.06.06.06.06.06.06.06.0
Words/Sentence5.85.56.07.15.65.87.65.16.3
Characters/word4.24.03.93.64.74.63.84.53.8
Readability
Passive sentences0%0%0%0%0%0%0%0%0%
Flesch Reading Ease70.380.785.591.359.275.285.06584.6
Flesch-Kincaid Grade Level4.83.32.82.26.44.23.35.43.0

APPENDIX 3: THE ORIGINAL SET OF TASKS

You will have to talk about the topic for 2 minutes. You have 1 minute to think about what you are going to say.

Tasks 1-5Tasks 6-9
1. Describe a city you have visited which has impressed you.
You should say:
Where it is situated
Why you visited it
What you liked about it
And explain why you prefer it to other cities.
6. Describe a teacher who has influenced you in your education.
You should say:
Where you met them
What subject they taught
What was special about them
And explain why this person influenced you so much.
2. Describe a competition (or contest) that you have entered.
You should say:
When the competition took place
What you had to do
How well you did it
And explain why you entered the competition (or contest).
7. Describe a film or a TV programme which has made a strong impression on you.
You should say:
What kind of film or TV programme it was, eg comedy
When you saw the film or TV programme
What the film or TV programme was about
And explain why this film or TV programme made such an impression on you.
3. Describe a part-time/holiday job that you have done.
You should say:
How you got the job
What the job involved
How long the job lasted
And explain why you think you did the job well or badly.
8. Describe a memorable event in your life.
You should say:
When the event took place
Where the event took place
What happened exactly
And why this event was memorable for you.
4. Describe a museum, exhibition or art gallery that you have visited.
You should say:
Where it is
What made you decide to go there
What you particularly remember about the place
And explain why you would or would not recommend it to your friend.
9. Describe something you own which is very important to you.
You should say:
Where you got it from
How long you have had it
What you use it for
And explain why it is so important to you.
5. Describe an enjoyable event that you experienced when you were at school.
You should say:
What the event was
When it happened
What was good about it
And explain why you particularly remember this event.

APPENDIX 4: THE FINAL SET OF TASKS

You will have to talk about the topic for 2 minutes. You have 1 minute to think about what you are going to say.

Tasks A-DTasks E-H
A. Describe a city you have visited which has impressed you.
You should say:
Where it is situated
Why you visited it
What you liked about it
And explain why you prefer it to other cities.
E. Describe a teacher who has influenced you in your education.
You should say:
Where you met them
What subject they taught
What was special about them
And explain why this person influenced you so much.
B. Describe a part-time/holiday job that you have done.
You should say:
How you got the job
What the job involved
How long the job lasted
And explain why you think you did the job well or badly.
F. Describe a film or a TV programme which made a strong impression on you.
You should say:
What kind of film or TV programme it was (eg comedy)
When you saw it
What it was about
And explain why it made such an impression on you.
C. Describe a sports event that you have been to or seen on TV.
You should say:
What it was
Why you wanted to see it
What was the most exciting or boring part
And explain why it was good or bad.
G. Describe a memorable event in your life.
You should say:
When the event took place
Where the event took place
What happened exactly
And why this event was memorable for you.
D. Describe an enjoyable event that you experienced when you were at school.
You should say:
What the event was
When it happened
What was good about it
And explain why you particularly remember this event.
H. Describe something you own which is very important to you.
You should say:
Where you got it from
How long you have had it
What you use it for
And explain why it is so important to you.

APPENDIX 5: SPSS ONE-WAY ANOVA OUTPUT

Multiple Comparisons

(I) TASK(J) TASKMean Difference (I-J)Std. ErrorSig.95% CI Lower Bound95% CI Upper Bound
Task ATask B.3622.227861.000-.35911.0835
Task ATask C-.0185.225701.000-.7330.6959
Task ATask D.3824.223681.000-.32561.0905
Task ATask E.4487.227861.000-.27261.1700
Task ATask F.6891.22786.079-.03221.4104
Task ATask G.9103*.22786.003.18901.6315
Task ATask H.7853*.22786.019.06401.5065
Task BTask A-.3622.227861.000-1.0835.3591
Task BTask C-.3807.227861.000-1.1020.3406
Task BTask D.0203.225861.000-.6947.7352
Task BTask E.0865.230001.000-.6415.8146
Task BTask F.3269.230001.000-.40111.0550
Task BTask G.5481.23000.507-.18001.2761
Task BTask H.4231.230001.000-.30501.1511
Task CTask A.0185.225701.000-.6959.7330
Task CTask B.3807.227861.000-.34061.1020
Task CTask D.4010.223681.000-.30711.1090
Task CTask E.4672.227861.000-.25401.1885
Task CTask F.7076.22786.061-.01371.4289
Task CTask G.9288*.22786.002.20751.6501
Task CTask H.8038*.22786.015.08251.5251
Task DTask A-.3824.223681.000-1.0905.3256
Task DTask B-.0203.225861.000-.7352.6947
Task DTask C-.4010.223681.000-1.1090.3071
Task DTask E.0663.225861.000-.6487.7812
Task DTask F.3067.225861.000-.40831.0216
Task DTask G.5278.22586.572-.18711.2428
Task DTask H.4028.225861.000-.31211.1178
Task ETask A-.4487.227861.000-1.1700.2726
Task ETask B-.0865.230001.000-.8146.6415
Task ETask C-.4672.227861.000-1.1885.2540
Task ETask D-.0663.225861.000-.7812.6487
Task ETask F.2404.230001.000-.4877.9684
Task ETask G.4615.230001.000-.26651.1896
Task ETask H.3365.230001.000-.39151.0646
Task FTask A-.6891.22786.079-1.4104.0322
Task FTask B-.3269.230001.000-1.0550.4011
Task FTask C-.7076.22786.061-1.4289.0137
Task FTask D-.3067.225861.000-1.0216.4083
Task FTask E-.2404.230001.000-.9684.4877
Task FTask G.2212.230001.000-.5069.9492
Task FTask H.0962.230001.000-.6319.8242
Task GTask A-.9103*.22786.003-1.6315-.1890
Task GTask B-.5481.23000.507-1.2761.1800
Task GTask C-.9288*.22786.002-1.6501-.2075
Task GTask D-.5278.22586.572-1.2428.1871
Task GTask E-.4615.230001.000-1.1896.2665
Task GTask F-.2212.230001.000-.9492.5069
Task GTask H-.1250.230001.000-.8531.6031
Task HTask A-.7853*.22786.019-1.5065-.0640
Task HTask B-.4231.230001.000-1.1511.3050
Task HTask C-.8038*.22786.015-1.5251-.0825
Task HTask D-.4028.225861.000-1.1178.3121
Task HTask E-.3365.230001.000-1.0646.3915
Task HTask F-.0962.230001.000-.8242.6319
Task HTask G.1250.230001.000-.6031.8531

APPENDIX 6: QUESTIONNAIRE ABOUT TASK 1

For each of the items below, circle the number that REFLECTS YOUR VIEWPOINT on a five point scale.

QuestionOption 1Option 2Option 3Option 4Option 5
1. The vocabulary in the task prompts was:Very easyVery difficult
1. The vocabulary in the task prompts was:12345
2. The grammatical structures in the task prompts were:Very easyVery difficult
2. The grammatical structures in the task prompts were:12345
3. Topic of the task was:Very familiarVery unfamiliar
3. Topic of the task was:12345
4. Information given in the task was:Very concreteVery abstract
4. Information given in the task was:12345
5. The planning time to complete (prepare for) the task was:Too longappropriateToo short
5. The planning time to complete (prepare for) the task was:12345
6. Time to complete the task was:Too longappropriateToo short
6. Time to complete the task was:12345
7. How much information did you use from the 4 short prompts provided in the task?1 = I used 100% of information provided in the task
7. How much information did you use from the 4 short prompts provided in the task?2 = I used 75% of information provided in the task
7. How much information did you use from the 4 short prompts provided in the task?3 = I used 50% of information provided in the task
7. How much information did you use from the 4 short prompts provided in the task?4 = I used 25% of information provided in the task
7. How much information did you use from the 4 short prompts provided in the task?5 = I did not use any information in the task at all
8. How did you use notes while you were speaking?1 = I read aloud my notes.
8. How did you use notes while you were speaking?2 = I referred to my notes line by line and looked up to speak.
8. How did you use notes while you were speaking?3 = I referred to my notes when I needed.
8. How did you use notes while you were speaking?4 = I prepared for my notes, but I did not use it.
8. How did you use notes while you were speaking?5 = I did not take my notes.

Thank you very much for your cooperation.

APPENDIX 7: QUESTIONNAIRE - UNCHANGED AND REDUCED TIME VERSIONS

For students responding to the unchanged versions and to the reduced response time versions

For each of the items below, circle the number that reflects your view point on the five point scale.

What I thought of or did before I started

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I read the task very carefully to understand what was required.12345
2. I thought of HOW to deliver my speech in order to respond well to the topic.12345
3. I thought of HOW to satisfy the audiences and examiners.12345
4. I understood the instructions for this speaking test completely.12345
5. I had ENOUGH ideas to speak about this topic.12345
6. I felt it was easy to produce enough ideas for the speech from memory.12345
7. I know A LOT about this type of speech, i.e., I know how to make a speech on this type of topic.12345
8. I know A LOT about other types of speaking test, e.g., interview, discussion.12345

What I thought of or did in planning stage

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I thought of MOST of my ideas for the speech BEFORE planning an outline.12345
2. During the period allowed for planning, I was conscious of the time.12345
3. I followed the 3 short prompts provided in the task when I was planning.12345
4. The information in the short prompts provided was necessary for me to complete the task.12345
5. I wrote down the points I wanted to make based on the 3 short prompts provided in the task.12345
6. I wrote down the words and expressions I needed to fulfil the task.12345
7. I wrote down the structures I need to fulfil the task.12345
8. I took notes only in ENGLISH.12345
9. I took notes only in my own language.12345
10. I took notes in both ENGLISH and own language.12345
11. I planned an outline on paper BEFORE starting to speak.1. Yes2. No
12. I planned an outline in my mind BEFORE starting to speak.1. Yes2. No
13. Ideas occurring to me at the beginning tended to be COMPLETE.12345
14. I was able to put my ideas or content in good order.12345
15. I practiced the speech in my mind WHILE I was planning.12345
16. After finishing my planning, I practiced what I was going to say in my mind until it was time to start.12345

What I thought of or did while I was speaking

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I felt it was easy to put ideas in good order.12345
2. I was able to express my ideas using suitable words.12345
3. I was able to express my ideas using correct grammar.12345
4. I thought of MOST of my ideas for the speech WHILE I was speaking.12345
5. WHILE I was speaking, I did not use some ideas that I had planned.12345
6. I was able to put sentences in logical order.12345
7. I was able to CONNECT my ideas smoothly in the whole speech.12345
8. I was conscious of the time WHILE I was making this speech.12345
9. I tried to finish speaking within the time.12345
10. I was listening and checking the correctness of the contents and their order WHILE I was making this speech.12345
11. I was listening and checking whether the contents and their order fit the topic WHILE I was making this speech.12345
12. I was listening and checking the correctness of sentences WHILE I was making this speech.12345
13. I was listening and checking whether the words fit the topic WHILE I was making this speech.12345
14. I felt it was easy to complete the task.12345
15. Comments on the above items:

Thank you for completing this questionnaire

APPENDIX 8: QUESTIONNAIRE - NO PLANNING VERSION

For students responding to the no planning versions

For each of the items below, circle the number that reflects your view point on the five point scale.

What I thought of or did before I started

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I read the task very carefully to understand what was required.12345
2. I thought of HOW to deliver my speech in order to respond well to the topic.12345
3. I thought of HOW to satisfy the audiences and examiners.12345
4. I understood the instructions for this speaking test completely.12345
5. I had ENOUGH ideas to speak about this topic.12345
6. I felt it was easy to produce enough ideas for the speech from memory.12345
7. I know A LOT about this type of speech, i.e., I know how to make a speech on this type of topic.12345
8. I know A LOT about other types of speaking test, e.g., interview, discussion.12345

What I thought of or did while I was speaking

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I felt it was easy to put ideas in good order.12345
2. I was able to express my ideas using suitable words.12345
3. I was able to express my ideas using correct grammar.12345
4. I thought of MOST of my ideas for the speech WHILE I was speaking.12345
5. WHILE I was speaking, I did not use some ideas that I had planned.12345
6. I was able to put sentences in logical order.12345
7. I was able to CONNECT my ideas smoothly in the whole speech.12345
8. I was conscious of the time WHILE I was making this speech.12345
9. I tried to finish speaking within the time.12345
10. I was listening and checking the correctness of the contents and their order WHILE I was making this speech.12345
11. I was listening and checking whether the contents and their order fit the topic WHILE I was making this speech.12345
12. I was listening and checking the correctness of sentences WHILE I was making this speech.12345
13. I was listening and checking whether the words fit the topic WHILE I was making this speech.12345
14. I felt it was easy to complete the task.12345
15. Comments on the above items:

Thank you for completing this questionnaire

APPENDIX 9: QUESTIONNAIRE - UNSCAFFOLDED VERSIONS

For students responding to the unscaffolded versions

For each of the items below, circle the number that reflects your view point on the five point scale.

What I thought of or did before I started

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I read the task very carefully to understand what was required.12345
2. I thought of HOW to deliver my speech in order to respond well to the topic.12345
3. I thought of HOW to satisfy the audiences and examiners.12345
4. I understood the instructions for this speaking test completely.12345
5. I had ENOUGH ideas to speak about this topic.12345
6. I felt it was easy to produce enough ideas for the speech from memory.12345
7. I know A LOT about this type of speech, i.e., I know how to make a speech on this type of topic.12345
8. I know A LOT about other types of speaking test, e.g., interview, discussion.12345

What I thought of or did in planning stage

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I thought of MOST of my ideas for the speech BEFORE planning an outline.12345
2. During the period allowed for planning, I was conscious of the time.12345
3. I followed the 3 short prompts provided in the task when I was planning.12345
4. The information in the short prompts provided was necessary for me to complete the task.12345
5. I wrote down the points I wanted to make based on the 3 short prompts provided in the task.12345
6. I wrote down the words and expressions I needed to fulfil the task.12345
7. I wrote down the structures I need to fulfil the task.12345
8. I took notes only in ENGLISH.12345
9. I took notes only in my own language.12345
10. I took notes in both ENGLISH and own language.12345
11. I planned an outline on paper BEFORE starting to speak.1. Yes2. No
12. I planned an outline in my mind BEFORE starting to speak.1. Yes2. No
13. Ideas occurring to me at the beginning tended to be COMPLETE.12345
14. I was able to put my ideas or content in good order.12345
15. I practiced the speech in my mind WHILE I was planning.12345
16. After finishing my planning, I practiced what I was going to say in my mind until it was time to start.12345

What I thought of or did while I was speaking

Item12345
Scale anchorstrongly disagreedisagreeno viewagreestrongly agree
1. I felt it was easy to put ideas in good order.12345
2. I was able to express my ideas using suitable words.12345
3. I was able to express my ideas using correct grammar.12345
4. I thought of MOST of my ideas for the speech WHILE I was speaking.12345
5. WHILE I was speaking, I did not use some ideas that I had planned.12345
6. I was able to put sentences in logical order.12345
7. I was able to CONNECT my ideas smoothly in the whole speech.12345
8. I was conscious of the time WHILE I was making this speech.12345
9. I tried to finish speaking within the time.12345
10. I was listening and checking the correctness of the contents and their order WHILE I was making this speech.12345
11. I was listening and checking whether the contents and their order fit the topic WHILE I was making this speech.12345
12. I was listening and checking the correctness of sentences WHILE I was making this speech.12345
13. I was listening and checking whether the words fit the topic WHILE I was making this speech.12345
14. I felt it was easy to complete the task.12345
15. Comments on the above items:

Thank you for completing this questionnaire

Author biodata

Cyril Weir

Cyril Weir has a PhD in language testing and has published widely in the fields of testing and evaluation. He is the author of Communicative Language Testing, Understanding and Developing Language Tests and Language Testing and Validation: an evidence based approach. He is the coauthor of Evaluation in ELT, An Empirical Investigation of the Componentiality of L2 Reading in English for Academic Purposes, Empirical Bases for Construct Validation: the College English Test - a case study, and Reading in a Second Language and co-editor of Continuity and Innovation: Revising the Cambridge Proficiency in English Examination 1913-2002. Cyril Weir has taught short courses, lectured and carried out consultancies in language testing, evaluation and curriculum renewal in over 50 countries worldwide. With Mike Milanovic of UCLES he is the series editor of the Studies in Language Testing series published by CUP and on the editorial board of Language Assessment Quarterly and Reading in a Foreign Language. Cyril Weir is currently Powdrill Professor in English Language Acquisition at the University of Bedfordshire, where he is also the Director of the Centre for Research in English Language Learning and Assessment (CRELLA) which was set up on his arrival in 2005.

Barry O’Sullivan

Barry O’Sullivan has a PhD in language testing, and is particularly interested in issues related to performance testing, test validation and test-data management and analysis. He has lectured for many years on various aspects of language testing, and is currently Director of the Centre for Language Assessment Research (CLARe) at Roehampton University, London. Barry’s publications have appeared in a number of international journals and he has presented his work at international conferences around the world. His book Issues in Business English Testing: the BEC revision project was published in 2006 by Cambridge University Press in the Studies in Language Testing series; and his next book is due to appear later this year. Barry is very active in language testing around the world and currently works with government ministries, universities and test developers in Europe, Asia, the Middle East and Central America. In addition to his work in the area of language testing, Barry taught in Ireland, England, Peru and Japan before taking up his current post.

Tomoko Horai

Tomoko Horai is a PhD student at Roehampton University, UK. She has an MA in Applied Linguistics and an MA in English Language Teaching, in addition to a MEd in TESOL/Applied Linguistics. She also has a number of years of teaching experience in a secondary school in Tokyo. Her current research interests are intra-task comparison and task difficulty in the testing of speaking. Her work has been presented at a number of international conferences including Language Testing Research Colloquium 2006, British Association of Applied Linguistics (BAAL) 2006, International Association of Teaching English as a Foreign Language (IATEFL) 2005 and 2006, Language Testing Forum 2005, and Japan Association of Language Teachers (JALT) 2004 and 2005.