Listening experiments
LISTENING EXPERIMENTS AND COMPARISONS
This site is under construction and every test is a pilot test. Choose how you would like to take part.
By entering you consent to your usage on this webiste being recorded without much guarantee.
Under construction
This site is under construction. Every test here is a pilot test: procedures, wording and results may change.
Psychoacoustics
Choose a listening test
METHOD PAIRWISE
Two sounds, one after the other; the listener says which. Ask that for every pair and a pile of "this one" answers turns into a proper scale, by the same arithmetic that ranks chess players from wins and losses. It never asks anyone for a number, which is what makes it immune to the habits that trouble the method above.
METHOD LIMITS
Change the sound in equal steps in one direction until the listener's answer changes, note where that happened, then do the same thing from the other direction. It is the oldest procedure in the field and the first one taught. The two directions rarely agree, and the gap between them is the second thing it measures.
METHOD ADJUSTMENT
The listener holds the control and moves it until two sounds match, or until something cancels out, then confirms. It is the most natural task here for someone who has never done a listening test, and it is fast, which is why the platform uses it wherever one sound has to be matched to another.
METHOD STAIRCASE
The sound follows the listener: a bit harder after each run of correct answers, a bit easier after a wrong one. Within a few dozen trials it is hovering around the hardest setting the listener can still manage. Most of the limits measured here are measured this way.
METHOD TRACKING
The listener holds a button down for as long as they can hear the sound and lets go when they cannot. Holding makes it quieter, letting go makes it louder again, and meanwhile the pitch drifts slowly upward. The zigzag they trace out is the answer: a whole curve, in four minutes, with no separate measuring step.
METHOD ORDERING
Every item is on one screen and the listener can switch between them freely, as often as they like. They drag them into order, or score each one, or pick just the best and the worst. Five ways of answering, one piece of arithmetic behind them, and it is the method behind almost every blind test of real audio here.
METHOD MAPPING
Two columns of items share one screen. The listener drags from the left to the right to say which things belong together. Either side can contain playable sounds or written descriptors, and each test decides whether the map is one-to-one, one-to-many or limited to a fixed number of connections.
METHOD FORCED CHOICE
One recording at a time, and a short row of labels for each question asked about it. The listener must pick one label per question — no slider, no "unsure" — so every clip ends up in exactly one box. A test can ask more than one question of the same clip, and can carry a reference answer, such as a classifier's verdict, to compare the listener against.
METHOD ANOMALY
One homogeneous sound runs continuously. The software introduces brief events into it near the edge of noticeability, at irregular times, and the listener has one main button: press the moment something actually stands out. The anomaly strength is adjusted separately for each condition, so the result is a curve of how strong the disturbance had to be before it broke through the carrier.
METHOD VIDEO
A short clip is played and the listener says what they got. What makes it a method of its own is that the picture and the sound need not come from the same recording: the platform takes a set of clips and plays one clip's face with another clip's voice. Ask what was heard, then ask the same clips what the mouth said, and the two answers between them say how much each sense moved the other.
Analysis by test method
| Test method | Primary analysis |
|---|---|
| LIMITS | |
The sound changes step by step, once from easy to hard and once from hard to easy. The analysis notes the point where the listener's answer changes in each direction. Repeating those journeys gives an average transition point for each direction. The gap between the two averages shows how much the answer depended on where the journey began. Together, the averages and their gap provide an understandable estimate of the listener's boundary and its stability. |
|
| ADJUSTMENT | |
The listener moves a control until two sounds seem to match or a chosen effect disappears. The analysis averages the final settings to find the listener's typical match. It also measures how widely those settings vary, which shows how consistently the listener can repeat the judgement. Results that began above the likely match are compared with results that began below it. Any difference between those starting directions warns that the control's initial position influenced the answer. |
|
| CONSTANT | |
The test presents a fixed set of sound levels many times in a shuffled order. The analysis calculates how often the listener gave the target answer at each level. Those percentages are joined with a smooth S-shaped curve showing how the answer changes as the sound becomes easier to hear. A chosen point on that curve, often where the listener is correct most but not all of the time, is reported as the threshold. The curve also shows whether the change from guessing to reliable hearing was sharp or gradual. |
|
| STAIRCASE | |
The test makes the next sound harder after successful answers and easier after a mistake. Each switch from getting harder to getting easier, or back again, is called a reversal. After the early settling-in period, the analysis averages the later reversal levels to estimate the listener's threshold. It also measures how spread out those reversals are. A tight cluster suggests a stable boundary, while a wide cluster says the estimate should be treated with more caution. |
|
| BAYES | |
This analysis starts with a broad range of plausible hearing curves and updates them after every answer. Curves that explain the listener's answers well become more likely, while poor explanations become less likely. The final result gives the most likely threshold and a range showing how uncertain that estimate remains. It also estimates how quickly performance improves as the sound becomes easier. A separate lapse estimate allows for occasional mistakes even when the sound should have been obvious. |
|
| TRACKING | |
The listener holds a button while a changing sound is audible and releases it when it disappears. This produces a moving trace rather than a set of separate right-or-wrong answers. The analysis first allows for the short delay between hearing a change and pressing or releasing the button. It then smooths small hand movements and momentary slips that are unlikely to reflect a real change in hearing. The midpoint between the upward and downward parts of the cleaned trace becomes the estimated hearing boundary over time or pitch. |
|
| CONFIDENCE | |
Some trials contain a signal and others contain only the background sound. A correct report of the signal is a hit, while reporting one when none was present is a false alarm. The analysis combines both rates so that a person who presses for everything does not appear unusually sensitive. One number describes how well the listener can separate signal from background. Another describes their response style, from cautious to willing to report even uncertain impressions. |
|
| DISCRIMINATION | |
The listener hears two or more choices and must select the one that is different or contains the target. The analysis begins with the proportion of correct choices. It removes the success expected from guessing, which depends on how many choices each trial offered. The result is converted to a common sensitivity scale so different forced-choice designs can be compared fairly. A higher value means the sounds were easier for the listener to tell apart. |
|
| SCALING | |
The listener expresses how strong a sound seems by choosing a number or matching it to another kind of magnitude. The analysis compares those answers with the physical amount of sound that was presented. It fits a curve showing how perceived strength grows as the physical signal grows. The curve can reveal compression, where large physical changes feel smaller, or expansion, where they feel larger. Differences between listeners or conditions are judged from the shape and position of that relationship. |
|
| PAIRWISE | |
The listener compares two items at a time and chooses the one that better fits the question. Across many pairings, each choice acts like a small win for one item over another. The analysis combines those wins into a single scale, much like a ranking built from match results. It allows for some listeners to use the task differently or to be more consistent than others. The final positions show the relative order and distance between items, with uncertainty where the evidence is thin. |
|
| ORDERING | |
The listener places several items in order, rates them, or identifies the best and worst. The analysis turns those responses into an overall position for every item. It checks whether an item's place on the screen or order of presentation changed the answers. It also allows for listeners to differ in how they use ratings or make rankings. The result shows which differences between items are well supported and which may simply reflect variation in people or presentation. |
|
| MAPPING | |
The listener connects items in one column to items in another. The analysis keeps every link as a directed pair rather than turning the screen into an order or a scale. When the study declares an answer key, the submitted links are compared with it using both precision and recall, so extra links and missed links both count. Without a key, the map itself is the result. Per-item and whole-screen limits define whether the task is one-to-one, one-to-many or many-to-many. |
|
| FORCED CHOICE | |
The listener hears one sound and must choose one label for each question asked about it. The analysis counts how often each label was chosen. When the test declares a reference answer for a sound, such as a classifier's verdict, it reports how often the listener agreed with it and, for labels with a natural order, how many steps above or below it the listener sat on average. When the labels follow a measured property of the sound, Kendall's rank correlation shows how closely the listener's ordering tracks the measurement. Answers given before the sound had finished playing are counted as rushed. |
|
| REPORTING | |
The listener continuously reports what they hear while a sound continues or changes. The analysis measures how long it takes before each reported experience first appears. It then measures how long each experience lasts and how often the listener switches between them. If the test ends while an experience is still continuing, that unfinished duration is kept rather than treated as a normal ending. Together these measures describe the timing and stability of perception, not just which answer occurred most often. |
|
| ANOMALY | |
A steady sound plays while brief changes sometimes appear at unpredictable moments. The analysis counts how many real changes the listener detected and how often they pressed when nothing changed. It measures the delay between each change and the listener's response. Results are also compared across the session to see whether attention or sensitivity drifted with time. Taken together, these measures separate reliable detection from quick but indiscriminate button pressing. |
|
| VIDEO | |
The listener watches a face and reports what they hear in its speech. Some trials pair matching sound and picture, while others deliberately pair a voice with a different mouth movement. The analysis compares the pattern of answers in those matching and conflicting conditions. A shift toward what the mouth appeared to say shows that vision influenced the heard answer. Comparing the opposite mismatch as well helps distinguish a genuine audiovisual effect from a simple preference for one response. |
|