WEBVTT 00:00:00.000 --> 00:00:06.620 The purpose of this seminar series is to talk about new developments in assessment, in educational assessment. 00:00:07.500 --> 00:00:13.980 And one of the things that we're very keen to do is to promote the idea that we can discuss complicated ideas 00:00:13.980 --> 00:00:22.640 and decide how we put them across to various audiences, and particularly to the public, and indeed using the press for that. 00:00:22.640 --> 00:00:42.140 So the topic of this evening's seminar is measurement error, and we want to think about the possibility of talking about that more frankly in terms of what the public understands by that concept, because that kind of openness and transparency is what society requires of us these days. 00:00:42.600 --> 00:00:50.100 So Paul's presentation will make us think about that, and I'm sure we'll have a really lively and interesting discussion at the end of it. 00:00:52.640 --> 00:01:02.040 I'm very pleased to introduce Paul Newton as our fourth speaker in the series this year. 00:01:02.980 --> 00:01:07.320 Paul has worked in various assessment organizations. 00:01:07.580 --> 00:01:11.300 I think he's going to tell us, so I won't steal that bit of his thunder. 00:01:11.900 --> 00:01:18.580 But Paul is very interested in the idea of how we explain ourselves 00:01:18.580 --> 00:01:23.800 and our work in assessment to the general public, to the non-specialist world. 00:01:24.440 --> 00:01:31.520 And the topic of his talk this evening, recognising the error of our ways, 00:01:31.840 --> 00:01:38.040 well, I think we'll begin to guess what it is he might be thinking of helping us to think about. 00:01:38.200 --> 00:01:41.200 So, Paul, welcome. Thank you very much for coming. 00:01:42.420 --> 00:01:46.720 Well, about 11 years ago, I was working as a researcher in King's College, that's London, 00:01:46.720 --> 00:01:59.148 and I was asked to give a presentation to the School of Education and I thought it would be a good idea to talk about exam standards I recently been working on exam standards at the AEB that the first organisation 00:01:59.888 --> 00:02:06.928 and Gillian Shepard had recently announced plans to rationalise the number of exam boards and the number of syllabuses as well, 00:02:07.328 --> 00:02:11.308 in response to worrying evidence that exam standards hadn't been maintained. 00:02:12.308 --> 00:02:15.468 So I decided to be fairly cavalier, as was my way, 00:02:15.548 --> 00:02:18.468 and to argue that Gillian Shepard and Channel 4 Dispatches, 00:02:18.608 --> 00:02:20.308 some of you will remember that programme, I'm sure, 00:02:20.768 --> 00:02:24.688 and various academics of the day constructed problems of comparability 00:02:24.688 --> 00:02:26.748 that didn't actually exist, 00:02:27.348 --> 00:02:30.308 particularly the idea that the boards were competing for market share 00:02:30.308 --> 00:02:32.208 by lowering their standards. 00:02:34.148 --> 00:02:37.588 By that time, I'd been influenced by Mike Creswell and Dylan William 00:02:37.588 --> 00:02:41.208 to think that public trust was really central to the maintenance of standards. 00:02:41.308 --> 00:02:46.908 And I was arguing that Gillian Shepard and co were threatening that trust quite unnecessarily. 00:02:48.128 --> 00:02:54.008 Why didn't they just trust the examining boards, good eggs that they are, to do their best in the best interests of learners? 00:02:55.848 --> 00:03:01.988 But by the end of the presentation, one particular eminent professor had become extremely irritated with me. 00:03:03.108 --> 00:03:07.028 Trust the examining boards? Why on earth should we trust the examining boards? 00:03:07.528 --> 00:03:09.848 They're nothing but merchants of secrets and lies. 00:03:09.848 --> 00:03:13.568 I'm paraphrasing that bit, I can't remember exactly what to say 00:03:13.568 --> 00:03:17.428 It might not have been far off though from what I remember 00:03:17.428 --> 00:03:21.628 I think I'd been a bit naive to say the least 00:03:21.628 --> 00:03:25.248 I genuinely had no idea of the extent to which many academics did 00:03:25.248 --> 00:03:28.648 and possibly still do distrust the exam boards 00:03:28.648 --> 00:03:33.408 and I think that her belief that the boards actively conspired to conceal the error 00:03:33.408 --> 00:03:35.368 of their ways genuinely surprised me 00:03:35.368 --> 00:03:39.448 because I'd just come from an awarding body research department and I actually thought 00:03:39.448 --> 00:03:41.528 that we were quite open and transparent. 00:03:42.308 --> 00:03:43.748 And the argument in a nutshell is this. 00:03:43.864 --> 00:03:48.604 It's that we need to be more open and transparent about error than we presently are. 00:03:49.304 --> 00:03:56.724 When I say we, I mean all of us, us being awarding bodies, regulators, departments, test developers and so on. 00:03:58.024 --> 00:04:03.484 Well it strikes me that the way in which we understand measurement error today can be traced directly back to this paper by Edgeworth. 00:04:04.404 --> 00:04:10.684 And his basic observation was that if you give a group of examiners the same script to Mark, then they're all different marks. 00:04:10.684 --> 00:04:17.864 So even though exactly the same script has been given to all of the examiners, there will inevitably be some kind of variability in the marks awarded. 00:04:18.324 --> 00:04:26.564 And he explained, whatever precautions have been taken to secure unity of standard, there will occur a certain divergence between the verdict of competent examiners. 00:04:27.124 --> 00:04:32.244 Say, four marks are 30, then if one examiner marks 20, another might mark 21, and another 19. 00:04:33.124 --> 00:04:39.124 If we tabulate the marks given by the different examiners, they would tend to be disposed after the fashion of a gendarme's hat. 00:04:39.124 --> 00:04:41.244 a jean d'armes hat 00:04:41.244 --> 00:04:45.264 no not that kind of jean d'armes hat 00:04:45.264 --> 00:04:47.004 this kind of jean d'armes hat 00:04:47.004 --> 00:04:48.944 because it is 1888 00:04:48.944 --> 00:04:52.104 we'd probably call it the bell curve nowadays 00:04:52.104 --> 00:04:54.884 anyhow 00:04:54.884 --> 00:05:01.044 okay so edgeworths with edging towards a fundamental point 00:05:01.044 --> 00:05:04.644 and the point was if you want to work out the degree of measurement error 00:05:04.644 --> 00:05:06.464 that's associated with an examination mark 00:05:06.464 --> 00:05:09.584 you need to start by working out what constitutes measurement truth. 00:05:10.784 --> 00:05:14.784 But if equally competent judges award different marks to the same script, 00:05:15.064 --> 00:05:18.184 how can you decide which mark is correct and which are incorrect? 00:05:19.784 --> 00:05:23.344 Well, according to Edgeworth, you start by drawing an analogy with physical measurement. 00:05:24.324 --> 00:05:29.284 And the analogy might be if one person measures a piece of string once with a single ruler, 00:05:29.664 --> 00:05:44.672 then they might accidentally get the measurement wrong But if a lot of people measure the same piece of string many times with many rulers then the number of times that they overestimate the length of the piece of string is likely to be balanced by the amount of times that they underestimate the length of the piece 00:05:44.672 --> 00:05:49.472 of string. And so the average length across all of those measurements will probably be a good 00:05:49.472 --> 00:05:55.452 estimate of the true length of the piece of string. So according to Edgeworth, this central figure, 00:05:55.692 --> 00:06:00.192 which is, or may be supposed to be, assigned by the greatest number of equally competent judges, 00:06:00.192 --> 00:06:03.752 is to be regarded as the true value of the Latin prose, 00:06:04.352 --> 00:06:06.152 just as the true weight of a body is determined 00:06:06.152 --> 00:06:08.772 by taking the mean of several discrepant measurements. 00:06:09.512 --> 00:06:09.632 Okay? 00:06:13.352 --> 00:06:15.352 And, of course, once the true mark has been defined, 00:06:15.452 --> 00:06:18.072 then you can take any deviation from it as an error. 00:06:18.712 --> 00:06:20.592 And he says, I think it's intelligible to speak 00:06:20.592 --> 00:06:23.112 of the mean judgment of competent critics as the true judgment, 00:06:23.592 --> 00:06:25.312 and deviations from that mean as errors. 00:06:26.812 --> 00:06:29.472 So Edgeworth developed this concept of measurement error 00:06:29.472 --> 00:06:31.572 that was able to deal with examinations quite well. 00:06:32.932 --> 00:06:35.452 He illustrated it in terms of the errors that lead to variability 00:06:35.452 --> 00:06:37.312 in the marks awarded by examiners, 00:06:37.592 --> 00:06:39.992 but he also acknowledged that there were lots of other errors. 00:06:40.612 --> 00:06:44.292 For example, the errors that arise from a candidate being out of sorts 00:06:44.292 --> 00:06:45.192 on the day of the examination, 00:06:45.652 --> 00:06:49.732 or the paper being specially adapted to the accidents of his taste or reading, 00:06:50.252 --> 00:06:54.112 or the circumstance that the questions are not fairly representative of the subject. 00:06:54.112 --> 00:06:57.772 And these are all direct quotations from his 1888 paper. 00:06:57.772 --> 00:07:00.712 so basically 120 years ago 00:07:00.712 --> 00:07:02.792 he developed a pretty sophisticated model 00:07:02.792 --> 00:07:05.172 of measurement error and it's one that we're still 00:07:05.172 --> 00:07:06.392 reliant upon today 00:07:06.392 --> 00:07:13.052 nowadays though we would tend to talk about 00:07:13.052 --> 00:07:15.092 measurement error under the heading of reliability 00:07:15.092 --> 00:07:17.392 and I like to think of reliability 00:07:17.392 --> 00:07:19.492 as quantifying the luck of the draw 00:07:19.492 --> 00:07:21.572 and that's where the principle of replication 00:07:21.572 --> 00:07:22.212 comes in 00:07:22.212 --> 00:07:25.492 principle of replication asks the hypothetical 00:07:25.492 --> 00:07:26.872 question like 00:07:26.872 --> 00:07:27.612 what 00:07:27.728 --> 00:07:30.788 what if a candidate happened to have sat the exam on a different day? 00:07:31.188 --> 00:07:34.768 Or what if the paper happened to have comprised a different set of questions? 00:07:35.308 --> 00:07:38.188 What if the script happened to have been marked by a different marker? 00:07:38.928 --> 00:07:41.608 Or what if a different awarding committee had been appointed? 00:07:42.828 --> 00:07:44.768 Would the same grave have been awarded? 00:07:44.928 --> 00:07:46.028 That's the fundamental question. 00:07:47.868 --> 00:07:50.568 These hypothetical replications, the ones on the slide, 00:07:50.568 --> 00:07:55.108 they're quite legitimate in the sense that the candidate might have sat the exam on a different day 00:07:55.108 --> 00:07:57.908 and the paper might have comprised a different set of questions, 00:07:58.028 --> 00:08:00.008 a different marker might have marked the script 00:08:00.008 --> 00:08:02.668 and a different awarding committee might have been appointed. 00:08:03.588 --> 00:08:08.908 So we're not trying to quantify the impact of maladministration or procedural error. 00:08:09.768 --> 00:08:13.368 We're trying to estimate the amount of error that inevitably remains 00:08:13.368 --> 00:08:18.208 even when robust examination procedures have been followed to the letter. 00:08:19.148 --> 00:08:21.888 So we're talking about measurement error, not procedural error 00:08:21.888 --> 00:08:22.968 and that's an important point. 00:08:23.808 --> 00:08:27.768 But what we really want to know is the combined effect of all sources of error 00:08:27.768 --> 00:08:30.628 on the overall reliability of results from national curriculum testing. 00:08:31.168 --> 00:08:33.728 And here it's probably fair to say that we know even less. 00:08:35.128 --> 00:08:39.328 In recent years, the closest we've come to a figure on the overall reliability of results 00:08:39.328 --> 00:08:42.668 from national curriculum testing is Dylan Williams' estimate of 30%. 00:08:42.668 --> 00:08:47.708 And the kind of key quote that's referred to a lot is on the overhead. 00:08:48.948 --> 00:08:51.888 Basically, what Dylan did is to develop a statistical simulation 00:08:51.888 --> 00:08:57.148 to model the likely impact of error on the overall reliability results, 00:08:57.468 --> 00:09:00.308 assuming different values for the reliability coefficient. 00:09:00.868 --> 00:09:04.748 So according to his simulation, with a coefficient of 0.80, 00:09:05.568 --> 00:09:09.928 a key stage 2 test will award incorrect levels to 32% of students. 00:09:11.128 --> 00:09:25.716 And with a coefficient of 0 27 of students will be misclassified So essentially for coefficients like this we looking at an error rate in the region of about 30 And this figure of 30 has been doing the round since 2001 00:09:26.876 --> 00:09:28.756 People do seem to be surprised by it, 00:09:28.936 --> 00:09:31.216 but no one's actually published a response to it, 00:09:31.376 --> 00:09:34.156 either to validate the claim or to invalidate it. 00:09:36.376 --> 00:09:39.516 Just recently, though, I've had the opportunity to work with the NFER 00:09:39.516 --> 00:09:43.416 to revisit some of their equating data for Key Stage 2 English. 00:09:44.236 --> 00:09:54.716 Now, their equating design required students to sit a pre-test version of the 2006 test just a few weeks before the live version of the 2005 test. 00:09:55.416 --> 00:10:00.356 And that enabled NFR to link standards from the 2005 version to the 2006 version. 00:10:01.456 --> 00:10:07.836 But it also enables a neat estimate of parallel form reliability, because you've got the same students sit in both tests. 00:10:08.336 --> 00:10:11.056 Parallel form reliability is very straightforward, really. 00:10:11.056 --> 00:10:14.656 you produce two different but interchangeable forms of the same test. 00:10:15.536 --> 00:10:18.136 You administer both of them to a single group of students 00:10:18.136 --> 00:10:19.476 within a short period of time. 00:10:20.516 --> 00:10:26.736 Then for each student, you simply compare the levels awarded by the two tests. 00:10:28.416 --> 00:10:30.656 If they get similar results on the two tests, 00:10:30.836 --> 00:10:32.976 then the tests are reliable. It's as simple as that. 00:10:34.056 --> 00:10:37.696 And this is probably as comprehensive an estimate of reliability as you can get. 00:10:37.696 --> 00:10:42.036 so when the NFER correlated marks across the two test forms 00:10:42.036 --> 00:10:44.016 that's 2005 versus 2006 00:10:44.016 --> 00:10:46.396 the coefficient of correlation was 0.85 00:10:46.396 --> 00:10:49.576 and apparently this figure can be taken as a good approximation 00:10:49.576 --> 00:10:50.976 of the reliability coefficient 00:10:50.976 --> 00:10:55.656 and remember when Dylan simulated a reliability coefficient of 0.85 00:10:55.656 --> 00:10:59.616 he calculated that 27% of students would be misclassified 00:10:59.616 --> 00:11:03.336 now interestingly when NFER compared levels 00:11:03.336 --> 00:11:06.696 across the 2005 and 2006 test versions 00:11:06.696 --> 00:11:11.476 they found that 73% of students were awarded the same level, i.e. 00:11:11.592 --> 00:11:14.632 exactly 27% received a different classification. 00:11:16.172 --> 00:11:18.332 But there's a very subtle but important difference here. 00:11:18.872 --> 00:11:21.272 Dylan estimated that 27% of students 00:11:21.272 --> 00:11:23.912 would be misclassified from a single administration, 00:11:25.472 --> 00:11:28.912 whereas NFER estimated that 27% would be differently classified 00:11:28.912 --> 00:11:30.812 from two test administrations. 00:11:33.332 --> 00:11:35.992 I'm going to have to use a rough analogy to make this point here 00:11:35.992 --> 00:11:37.052 because it's quite complicated. 00:11:39.252 --> 00:11:41.012 Imagine you've got a pass-fail test, 00:11:41.012 --> 00:11:44.652 which disagrees on the classification of 30% of students. 00:11:45.652 --> 00:11:50.172 And you've got that information from a parallel forms reliability study 00:11:50.172 --> 00:11:52.632 and you've counted how many students received different levels. 00:11:53.912 --> 00:11:55.912 So having administered the test twice, 00:11:56.012 --> 00:11:59.632 you're now uncertain about the classification of 30% of students. 00:12:01.712 --> 00:12:04.792 However, if you'd only administered one of those test forms, 00:12:05.412 --> 00:12:07.592 half of those students of uncertain levels, 00:12:07.592 --> 00:12:08.972 so half of that 30%, 00:12:08.972 --> 00:12:11.652 would have received the correct classification by chance alone, 00:12:11.992 --> 00:12:13.772 because it's a pass-fail test after all. 00:12:15.052 --> 00:12:19.052 That leads me to think that classification accuracy from one test administration 00:12:19.052 --> 00:12:22.652 is going to be higher than classification consistency from two. 00:12:24.892 --> 00:12:28.452 And from some very rough calculations that I did, very rough indeed, 00:12:28.952 --> 00:12:35.192 I estimated that 73% classification consistency from NPR's data 00:12:35.192 --> 00:12:39.872 probably translates into a figure of around 84% classification accuracy. 00:12:40.472 --> 00:12:45.672 And that's 16% misclassification, not the 27% that Dylan arrived at. 00:12:46.872 --> 00:12:52.192 So, which of these figures, 16%, 27% is more plausible? 00:12:52.852 --> 00:12:55.032 Well, to be honest, I'm not entirely sure. 00:12:56.372 --> 00:12:57.932 I'm sure about two things, though. 00:12:57.992 --> 00:13:14.700 The first is that we need to do more work in this area second is whichever estimate of error you go for there probably going to be their own margin of error around them and probably quite a large one Dylan Williams figure of 30 has been doing the round since 2001 People 00:13:14.700 --> 00:13:19.760 do seem to be surprised by it. Agencies like QCA have found it hard to know how to respond 00:13:19.760 --> 00:13:24.800 to it. Last year, during the Select Committee Inquiry and Testing and Assessment, Annette 00:13:24.800 --> 00:13:31.560 Brooke challenged Ken Boston to comment on Dylan's 30%. And Ken responded, well, error exists. 00:13:31.680 --> 00:13:36.460 As I said before, this is a process of judgment. Error exists, and error needs to be identified and 00:13:36.460 --> 00:13:41.820 rectified where it occurs. I am surprised at the figure of 30%. We've been looking at the range of 00:13:41.820 --> 00:13:46.960 tests and examinations for some time. We think this is a very high figure, but whatever it is, 00:13:46.960 --> 00:13:53.520 it needs to be capable of being identified and corrected. I can take that thought experiment a 00:13:53.520 --> 00:13:58.160 stage further, if you like, by asking the question, could it ever be legitimate for the 00:13:58.160 --> 00:14:02.820 head of a regulatory authority to stand up on results day and say, well, we congratulate 00:14:02.820 --> 00:14:07.120 students on their excellent exam results, but some of you haven't done quite as well 00:14:07.120 --> 00:14:13.100 as you think, and some of you have actually done better. Because establishing the size 00:14:13.100 --> 00:14:16.760 of the error rate is challenging enough, but knowing how to convey it appropriately is 00:14:16.760 --> 00:14:18.420 another thing, again. 00:14:18.420 --> 00:14:28.280 Okay, well, as I learned from an eminent professor of education back in 1997, the examining boards have had a reputation for being secretive. 00:14:29.160 --> 00:14:33.760 In fact, there was a time when very little information indeed escaped the four walls of examining board buildings. 00:14:34.720 --> 00:14:39.140 Error was understood within those four walls, wasn't necessarily spoken of beyond them. 00:14:40.040 --> 00:14:45.000 And the academic Stephen Wiseman wrote in 1961, I should point out, for those of you who can't quite see the date. 00:14:45.000 --> 00:14:49.000 boards seem to have strong objections 00:14:49.000 --> 00:14:51.220 to revealing the mysteries to outsiders 00:14:51.220 --> 00:14:53.980 there have undoubtedly been cases of inquiries 00:14:53.980 --> 00:14:55.340 where publication would have been 00:14:55.456 --> 00:14:59.336 in the interest of education and would have helped to prevent the spread of horror stories 00:14:59.336 --> 00:15:05.116 about such things as lack of equivalence, which is an inevitable concomitant of the present cloak of secrecy. 00:15:06.316 --> 00:15:10.916 In effect, he was claiming that being closed and opaque was fueling rumour and suspicion 00:15:10.916 --> 00:15:15.696 and that greater openness and transparency could actually help to quash those rumours and suspicions. 00:15:17.316 --> 00:15:22.416 But of course, greater openness and transparency runs the risk of owning up to the possibility, 00:15:22.416 --> 00:15:25.236 in fact, the inevitability of measurement error. 00:15:25.456 --> 00:15:29.036 So the question is, were the boards brave enough to step up to that challenge? 00:15:30.936 --> 00:15:32.176 Unfortunately, the answer is yes. 00:15:34.536 --> 00:15:36.896 Because between the early 60s and the mid-70s, 00:15:36.976 --> 00:15:41.036 they conducted a series of investigations into inter-board comparability of standards, 00:15:41.296 --> 00:15:43.896 and they published a compendium of results in 1978. 00:15:44.556 --> 00:15:46.716 And the preface to that compendium read, 00:15:47.216 --> 00:15:49.196 In presenting this booklet to the public, 00:15:49.736 --> 00:15:52.196 we in the GCE boards have found ourselves in a dilemma. 00:15:52.196 --> 00:15:59.556 If we merely state that comparability exercises are regularly conducted and do not show our hand, we appear to have something to hide. 00:16:00.276 --> 00:16:05.196 If we try to explain them, their complexities and limitations invite misunderstanding and misrepresentation. 00:16:06.236 --> 00:16:09.876 On balance, the preferable alternative seem to be to publish and be damned. 00:16:10.436 --> 00:16:11.476 We have and probably shall be. 00:16:13.776 --> 00:16:18.696 Well, I think that this ushered in a new dawn of openness for the examining boards. 00:16:19.276 --> 00:16:21.796 And as their research departments have increased in strength, 00:16:21.896 --> 00:16:23.636 so too have their publication records. 00:16:24.636 --> 00:16:28.156 And I think nowadays these publications provide a public window 00:16:28.156 --> 00:16:31.916 into the world of examinations through which their limitations are laid bare. 00:16:32.836 --> 00:16:35.156 Kind of like a Damien Hirst cow in formaldehyde. 00:16:35.556 --> 00:16:37.416 It's not pretty, it's slightly shocking. 00:16:38.076 --> 00:16:40.056 I'm not even sure whether we want to look at it. 00:16:41.056 --> 00:16:54.571 But it there for everyone to see and the world hasn collapsed as a consequence OK as you can see from the side there some great examples of exam board research out there nowadays and these 00:16:54.571 --> 00:16:58.931 are some recent examples of research into reliability. There's a lot of good work going 00:16:58.931 --> 00:17:04.231 on and just as importantly, it's getting published. It's not always pretty. Typically when you 00:17:04.231 --> 00:17:09.451 do exams research, you're explicitly addressing the problem of error. Why isn't my assessment 00:17:09.451 --> 00:17:13.771 perfect, how can I make it better? So it's not going to be pretty and sometimes it might 00:17:13.771 --> 00:17:20.051 be a bit shocking. But it's crucial all the same. Because without research like this, 00:17:20.211 --> 00:17:24.071 we wouldn't be able to improve our exams. More to the point, without publishing this 00:17:24.071 --> 00:17:28.071 research, we wouldn't be able to persuade the public that we're committed to their improvement. 00:17:28.851 --> 00:17:32.811 The point is, we've got some really interesting research going on there. It's highbrow stuff. 00:17:33.031 --> 00:17:37.971 It's about identifying the phenomena of assessment and seeking explanations for them. 00:17:38.971 --> 00:17:42.471 Assumption being that if we can genuinely understand these phenomena of assessment, 00:17:42.931 --> 00:17:45.531 then we'll be able to predict and control them better. 00:17:46.171 --> 00:17:49.251 And that's really an aspiration for a proper science of educational assessment. 00:17:51.011 --> 00:17:54.631 But I think there are some aspects of research that we're not doing so thoroughly. 00:17:55.191 --> 00:17:57.811 Or at least if we are, we're not publishing them with such fervour. 00:17:58.731 --> 00:18:02.051 And these are the more mundane aspects, if you like, the monitoring research. 00:18:02.051 --> 00:18:05.831 as QCA and now as Ofqual 00:18:05.831 --> 00:18:08.411 we monitor comparability routinely 00:18:08.411 --> 00:18:10.411 we've got a rolling programme of research 00:18:10.411 --> 00:18:12.751 and the boards also monitor comparability 00:18:12.751 --> 00:18:15.091 and have done so for decades and decades 00:18:15.091 --> 00:18:17.091 but not reliability 00:18:17.091 --> 00:18:22.231 I don't think we know enough about the reliability of our tests and examinations 00:18:22.231 --> 00:18:25.031 we know they're not perfectly reliable 00:18:25.031 --> 00:18:28.011 but do we know exactly how reliable they are 00:18:28.011 --> 00:18:31.291 do we know whether they're more reliable now 00:18:31.291 --> 00:18:34.971 than they were 20 years ago, or 120 years ago, come to that. 00:18:36.111 --> 00:18:39.191 Do we know whether new examination structures will be more or less reliable