Sample stimuli

sample 0 sample 1 sample 2 sample 3 sample 4 sample 5 sample 6 sample 7 sample 8 sample 9

How to use

from brainscore_vision import load_benchmark
benchmark = load_benchmark("Rajalingham2018-i2n")
score = benchmark(my_model)

Model scores

Min Alignment Max Alignment

Rank

Model

Score

1
.664
2
.652
3
.646
4
.625
5
.624
6
.623
7
.617
8
.608
9
.607
10
.607
11
.606
12
.600
13
.600
14
.596
15
.594
16
.593
17
.590
18
.588
19
.586
20
.585
21
.585
22
.584
23
.584
24
.584
25
.582
26
.581
27
.579
28
.579
29
.578
30
.578
31
.578
32
.578
33
.577
34
.576
35
.575
36
.574
37
.573
38
.573
39
.573
40
.573
41
.572
42
.570
43
.570
44
.568
45
.566
46
.564
47
.564
48
.564
49
.563
50
.563
51
.562
52
.561
53
.561
54
.561
55
.560
56
.560
57
.560
58
.560
59
.558
60
.558
61
.555
62
.555
63
.555
64
.555
65
.555
66
.554
67
.554
68
.554
69
.552
70
.551
71
.549
72
.549
73
.549
74
.549
75
.549
76
.549
77
.546
78
.546
79
.546
80
.545
81
.545
82
.545
83
.545
84
.543
85
.543
86
.542
87
.541
88
.541
89
.541
90
.541
91
.540
92
.540
93
.539
94
.538
95
.537
96
.537
97
.537
98
.537
99
.536
100
.536
101
.536
102
.535
103
.535
104
.534
105
.534
106
.534
107
.534
108
.534
109
.533
110
.533
111
.532
112
.532
113
.531
114
.530
115
.528
116
.528
117
.528
118
.528
119
.528
120
.528
121
.527
122
.527
123
.527
124
.526
125
.526
126
.526
127
.524
128
.524
129
.524
130
.523
131
.523
132
.523
133
.523
134
.522
135
.522
136
.521
137
.521
138
.521
139
.521
140
.521
141
.521
142
.520
143
.520
144
.520
145
.520
146
.519
147
.518
148
.518
149
.517
150
.517
151
.516
152
.515
153
.515
154
.515
155
.515
156
.515
157
.514
158
.513
159
.513
160
.513
161
.512
162
.512
163
.512
164
.512
165
.511
166
.511
167
.511
168
.511
169
.511
170
.510
171
.509
172
.509
173
.508
174
.508
175
.507
176
.507
177
.506
178
.505
179
.505
180
.504
181
.503
182
.503
183
.503
184
.503
185
.503
186
.502
187
.502
188
.502
189
.500
190
.500
191
.500
192
.500
193
.499
194
.499
195
.499
196
.499
197
.499
198
.499
199
.498
200
.498
201
.497
202
.496
203
.496
204
.495
205
.494
206
.494
207
.493
208
.493
209
.492
210
.491
211
.490
212
.488
213
.488
214
.488
215
.488
216
.487
217
.487
218
.485
219
.484
220
.481
221
.481
222
.480
223
.480
224
.479
225
.479
226
.478
227
.478
228
.478
229
.477
230
.477
231
.477
232
.476
233
.475
234
.475
235
.475
236
.474
237
.474
238
.474
239
.474
240
.473
241
.472
242
.472
243
.471
244
.470
245
.470
246
.469
247
.466
248
.465
249
.464
250
.462
251
.461
252
.461
253
.458
254
.458
255
.456
256
.456
257
.454
258
.454
259
.452
260
.451
261
.451
262
.450
263
.449
264
.449
265
.448
266
.448
267
.448
268
.448
269
.447
270
.447
271
.447
272
.446
273
.446
274
.445
275
.445
276
.445
277
.444
278
.443
279
.443
280
.441
281
.440
282
.438
283
.438
284
.437
285
.437
286
.437
287
.435
288
.435
289
.434
290
.433
291
.433
292
.430
293
.428
294
.428
295
.427
296
.426
297
.425
298
.425
299
.424
300
.424
301
.419
302
.415
303
.413
304
.413
305
.410
306
.410
307
.410
308
.408
309
.407
310
.406
311
.405
312
.403
313
.401
314
.396
315
.395
316
.392
317
.386
318
.383
319
.381
320
.376
321
.375
322
.373
323
.372
324
.371
325
.370
326
.370
327
.370
328
.370
329
.367
330
.366
331
.365
332
.363
333
.362
334
.360
335
.360
336
.358
337
.356
338
.354
339
.351
340
.348
341
.348
342
.346
343
.344
344
.341
345
.341
346
.335
347
.334
348
.333
349
.333
350
.333
351
.332
352
.330
353
.324
354
.324
355
.322
356
.320
357
.315
358
.311
359
.310
360
.307
361
.307
362
.306
363
.305
364
.292
365
.292
366
.291
367
.286
368
.286
369
.285
370
.284
371
.283
372
.279
373
.276
374
.276
375
.272
376
.270
377
.270
378
.267
379
.265
380
.263
381
.261
382
.259
383
.256
384
.256
385
.256
386
.256
387
.256
388
.256
389
.256
390
.256
391
.256
392
.255
393
.254
394
.251
395
.250
396
.245
397
.244
398
.243
399
.243
400
.242
401
.234
402
.231
403
.226
404
.225
405
.220
406
.219
407
.216
408
.211
409
.211
410
.209
411
.209
412
.208
413
.200
414
.187
415
.186
416
.177
417
.167
418
.167
419
.165
420
.161
421
.160
422
.157
423
.157
424
.156
425
.150
426
.148
427
.144
428
.137
429
.131
430
.129
431
.127
432
.119
433
.116
434
.114
435
.113
436
.112
437
.108
438
.108
439
.107
440
.104
441
.104
442
.103
443
.103
444
.102
445
.101
446
.098
447
.096
448
.095
449
.092
450
.090
451
.084
452
.084
453
.083
454
.083
455
.082
456
.078
457
.076
458
.075
459
.071
460
.071
461
.067
462
.065
463
.065
464
.061
465
.060
466
.060
467
.057
468
.054
469
.054
470
.049
471
.047
472
.046
473
.046
474
.045
475
.045
476
.043
477
.041
478
.040
479
.032
480
.030
481
.027
482
.023
483
.020
484
.020
485
.014
486
.014
487
.012
488
.011
489
.011
490
.010
491
.009
492
.009
493
.009
494
.004
495
.000
496
.000
497
.000
498
.000
499
.000
500
.000
501
.000
502
.000
503
.000
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520

Benchmark bibtex

@article {Rajalingham240614,
                author = {Rajalingham, Rishi and Issa, Elias B. and Bashivan, Pouya and Kar, Kohitij and Schmidt, Kailyn and DiCarlo, James J.},
                title = {Large-scale, high-resolution comparison of the core visual object recognition behavior of humans, monkeys, and state-of-the-art deep artificial neural networks},
                elocation-id = {240614},
                year = {2018},
                doi = {10.1101/240614},
                publisher = {Cold Spring Harbor Laboratory},
                abstract = {Primates{	extemdash}including humans{	extemdash}can typically recognize objects in visual images at a glance even in the face of naturally occurring identity-preserving image transformations (e.g. changes in viewpoint). A primary neuroscience goal is to uncover neuron-level mechanistic models that quantitatively explain this behavior by predicting primate performance for each and every image. Here, we applied this stringent behavioral prediction test to the leading mechanistic models of primate vision (specifically, deep, convolutional, artificial neural networks; ANNs) by directly comparing their behavioral signatures against those of humans and rhesus macaque monkeys. Using high-throughput data collection systems for human and monkey psychophysics, we collected over one million behavioral trials for 2400 images over 276 binary object discrimination tasks. Consistent with previous work, we observed that state-of-the-art deep, feed-forward convolutional ANNs trained for visual categorization (termed DCNNIC models) accurately predicted primate patterns of object-level confusion. However, when we examined behavioral performance for individual images within each object discrimination task, we found that all tested DCNNIC models were significantly non-predictive of primate performance, and that this prediction failure was not accounted for by simple image attributes, nor rescued by simple model modifications. These results show that current DCNNIC models cannot account for the image-level behavioral patterns of primates, and that new ANN models are needed to more precisely capture the neural mechanisms underlying primate object vision. To this end, large-scale, high-resolution primate behavioral benchmarks{	extemdash}such as those obtained here{	extemdash}could serve as direct guides for discovering such models.SIGNIFICANCE STATEMENT Recently, specific feed-forward deep convolutional artificial neural networks (ANNs) models have dramatically advanced our quantitative understanding of the neural mechanisms underlying primate core object recognition. In this work, we tested the limits of those ANNs by systematically comparing the behavioral responses of these models with the behavioral responses of humans and monkeys, at the resolution of individual images. Using these high-resolution metrics, we found that all tested ANN models significantly diverged from primate behavior. Going forward, these high-resolution, large-scale primate behavioral benchmarks could serve as direct guides for discovering better ANN models of the primate visual system.},
                URL = {https://www.biorxiv.org/content/early/2018/02/12/240614},
                eprint = {https://www.biorxiv.org/content/early/2018/02/12/240614.full.pdf},
                journal = {bioRxiv}
            }

Ceiling

0.48.

Note that scores are relative to this ceiling.

Data: Rajalingham2018

240 stimuli match-to-sample task

Metric: i2n