
==== Front
Forensic Sci Int Synerg
Forensic Sci Int Synerg
Forensic Science International: Synergy
2589-871X
Elsevier

S2589-871X(24)00094-9
10.1016/j.fsisyn.2024.100547
100547
General
Training on dirty labels: Rejoinder to Kotsoglou and Biedermann
Asonov Dmitri
Sber Innovation and Research, Sberbank of Russia, Moscow, 117997, Russian Federation
Krylov Maksim MAKrylov@sberbank.ru
⁎
Ryabikina Anastasiya
Mikhailov Maksim
Internal Security Department, Sberbank of Russia, Moscow, 117997, Russian Federation
⁎ Corresponding author. MAKrylov@sberbank.ru
15 8 2024
2024
15 8 2024
9 100547© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keywords

Polygraph screening
Machine learning
Research methodology
==== Body
pmcMajor critique of our paper [1] by Kotsoglou et al. [2] is that we train on dirty (non-ground truth) labels. They state: “Limiting the scope of ML to mimicking the process of reaching conclusions by polygraph interviewers is tantamount to further enforcing and imposing current PS practice in its imperfect state of development and operation.” We respectfully clarify that we did exactly the opposite. We recognized that dirty labels will negatively affect the quality of models, i.e. examiner biases and mistakes may propagate into the automated scenario. To counter this propagation, and before going to production, we unleashed our results – the examiner error detection method – on the field train set with dirty labels. As stated in our paper, we found 30 problematic examiner conclusions in our train set of 2094 screenings. We then trained the production model with these 30 errors corrected.

In general, the problem of field labels contaminated by human expert errors is gaining traction, and is recently called “inversion problem” by Kleinberg et al. [3]. While to the best knowledge of Kleinberg et al. no technique exists yet to solve this problem, we consider our results a practical demonstration of a partial solution to the inversion problem.

We now respond to what we classify as auxiliary critique, citing the critique first.

“Asonov et al. argue that “screening errors are not only due to the method, but also due to human (polygraph examiner) errors”. We question whether this distinction is helpful in practice, as the polygraph interviewer is an integral part of the screening process and it is difficult to separate the two. … The lack of knowledge of the ground truth means that errors, i.e. a discrepancy between the examiner's conclusion and the ground truth, cannot be detected.”

It is easy to recognize most examiner errors in the notion of the error that we described in the original work. Examiner errors occur, for example, when an examiner is inexperienced, exhausted or distracted, or biased. These errors are detected when a professional, unbiased examiner reviews the screening. We agree that there is a small fraction of mistakes that require a high level of human expertise and attention to detect. We also agree that if the examiner error is defined as the mismatch between an unbiased, careful and professional examiner conclusion and ground truth then such errors are impossible to detect without knowledge of ground truth.

“… what Asonov et al. presumably refer to as the “method” is the procedural component that involves physical measurements of the interrogated person.”

Polygraph method is a common term in polygraph practice and scientific research, referring to the questioning tactics and decision-making rules that examiners must follow. Examples include CQT (Control Question Technique) and CIT (Concealed Information Techniques), with many specific decision-making rules attached. Our original paper references overviews of various methods in the sentence “There are many good overviews of classical polygraph and questioning methods […].”

"In this paper we critically examine a recent proposal to apply ML to polygraph screening results”

Our ML-based second opinion is possibly the first one on the record that was thoroughly tested, accepted, and is now being routinely used by examiners in the field. But the proposal to apply ML to polygraph examinations is far from recent; first approach made by Ref. [4] in 2002, and at least 48 papers on similar topics, including ours, have been published since then [5].

We note that some auxiliary critique is not directly related to our results, but we chose to respond.

“… the field did not survive the 1920s because it had too many internal methodological inconsistencies and relied on too many idealisations.”

There is an ambiguity as to what Kotsoglou et al. meant by “the field”. If colleagues meant scientific research field in the area of lie detection and polygraph techniques, then we strongly disagree. The research field is booming, with approximately 400 papers published since 2023 alone [6], the field does not look dead since 1920s at all.

“… such research can inappropriately legitimize otherwise scientifically invalid, indeed pseudo-scientific methods such as polygraph-based deception detection, especially when presented in a reputable scientific journal.”

Polygraph-based deception detection is a legitimate technique in many countries, including in the UK (where Kotsoglou works). Moreover, Kotsoglou et al. themselves state that recently, in “several jurisdictions, including England and Wales, [there] has been the increased use of polygraph-based interviewing techniques [by the criminal justice system]”. There is simply no need to legitimize it. Furthermore, calling the methods used by special government agencies for decades “pseudo-scientific” borders on calling these agencies unprofessional. According to this logic of Kotsoglou et al., which we do not share, journals that publish research where these methods are employed should be considered pseudo-scientific journals too, including the very journal where Kotsoglou et al. published their critique [7]. We would describe methods like polygraph-based deception detection as controversial (provoking strongly opposing viewpoints), high-stakes, and yet not a fully investigated phenomenon, but not “pseudo-scientific” in any way. This phenomenon is rooted inside even more intriguing and under-investigated phenomena of lie and truth, which partially explains the difficulties in the scientific investigations.

Kotsoglou et al. cannot or wish not conduct research that minimizes deficiencies of the widespread, de-facto, and in many cases de-jure, standard of probabilistic deception detection (polygraph) employed by police, special government agencies, and the private sector. However, what is the point then in spending time criticizing those who do? Instead, wouldn't it be more productive to research an alternative method free of deficiencies?

We agree with Kotsoglou et al. in that the historical, marketing name “lie detector” is a curse that has been reaping its fruits until these days. If it had been named “probabilistic lie detection using sensors and an expert”, we might have had fewer debates over the emotive “pseudo” word, and more energy would have been spent on the investigation of the phenomenon further. We also agree that classical polygraph allows to see a response to stimuli, rather than to detect a lie. The complex job of an examiner is to interpret the response, and where needed, to perform additional, more narrowed tests. Furthermore, we agree that polygraph test conclusions simply cannot be used as a sole source of information in internal or criminal investigations, partially because the method is prone to errors, but this is a point of view of all practitioners anyway.

Summarizing this part, we would like to cite our original work: “Our results neither justify nor solidify the practice of classical polygraph screenings. Rather, we consider our results as a temporary and partial patch that helps to eliminate a specific type of error of this method, until better methods are devised and put into practice. More broadly, we believe we make a step towards rethinking classical polygraph practices.”

“Polygraph interviewers do not directly “detect” lies, strictly speaking, but only indicate deception, as the terms DI and NDI imply. … In practice, however, this subtlety is often ignored because consumers of polygraph interview results often confuse indications of deception with lie detection.”

Here we fully concur. We already stated in our original work that standardization of examiner education has great potential. Similarly, educating amateur consumers is of paramount importance too.

In general, we see several problems in the decades-long struggle between the pro- and anti-polygraph camps. (i) There is still no single, clean, agreed upon list of agreements and disagreements between the camps. Many papers are a dispute between one author and another. Perhaps some neutral party, such as a journal, could take the initiative to mediate and compile such a list, involving as many scientists and reputable organizations as possible. This could be done, for example, through requests for voluntary admissions from both sides. For example, above, we had no problem admitting to many deficiencies of the polygraph highlighted in Ref. [2]. (ii) There is a third, neutral camp, that of police, federal agencies and corporate security, which is neither for nor against the polygraph. This camp of professional practitioners is well aware of the pitfalls of the classical polygraph methods and does its best to mitigate them while doing its job in fighting crime or corporate fraud. The neutral camp values the polygraph because they see examples where the polygraph helped to solve or prevent a serious crime, firsthand. This camp will switch to another, better tool if and when it emerges. The problem with this camp is that it rarely participates in any public discussions at all. (iii) The pro-polygraph and neutral camps each have their own internal disagreements on specific polygraph methods; neither camp is united. Moreover, although the neutral camp is much larger than the pro- and anti-polygraph camps combined, in most cases these people are not allowed to express their opinions publicly. Therefore, there is a cognitive trap: the anti-polygraph camp is a minority, but it is united and vocal, making it appear as if it is a majority.

We believe that getting all three camps around the table and producing a joint document is what was needed yesterday to start rethinking the classical polygraph.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

This paper reflects the authors' views only. The Bank is not responsible for any use that may be made of the information contained.
==== Refs
References

1 Asonov D. Krylov M. Omelyusik V. Ryabikina A. Litvinov E. Mitrofanov M. Mikhailov M. Efimov A. Building a second-opinion tool for classical polygraph Sci. Rep. 13 2023
2 Kotsoglou K. Biedermann A. Polygraph-based deception detection and machine learning. Combining the worst of both worlds? Forensic Sci. Int.: Synergy 9 2024
3 Kleinberg J. Ludwig J. Mullainathan S. Raghavan M. The inversion problem: why algorithms should infer mental state and not just predict behavior Perspect. Psychol. Sci. 2023
4 Slavkovic A. Evaluating Polygraph Data 2002 Carnegie Mellon University Available: https://www.stat.cmu.edu/tr/tr766/tr766.pdf
5 Rad D. Paraschiv N. Kiss C. Neural network applications in polygraph scoring—a scoping review Information 14 2023
6 Google Scholar Search for "lie Detection" and "polygraph" 2024 Available: https://scholar.google.com/scholar?as_ylo=2023&q=%22lie+detection%22%2C+%22polygraph%22
7 Ayoub A. Amjid M. A polygraph case study of sodomy and murder case Forensic Sci. Int.: Synergy 6 2023
