Contents type: Verbal, numerical, spatial. Period: 1995-2004
| 0 | ** |
| 1 | *** |
| 1.5 | * |
| 2 | ******** |
| 3 | ******** |
| 3.5 | * |
| 4 | ********* |
| 4.5 | ** |
| 5 | *** |
| 6 | *** |
| 7 | ****** |
| 8 | *** |
| 8.5 | ** |
| 9 | **** |
| 10 | * |
| 11 | *** |
| 12 | ** |
| 12.5 | ** |
| 13 | * |
| 14 | * |
| 15 | ** |
| 16 | * |
| 16.5 | * |
| 17 | * |
| 21.5 | * |
| 25 | * |
| 27 | * |
n = 66
| 0 | * |
| 1 | * |
| 1.5 | * |
| 2 | ****** |
| 3 | ******** |
| 3.5 | * |
| 4 | ******* |
| 4.5 | ** |
| 5 | *** |
| 6 | *** |
| 7 | ****** |
| 8 | *** |
| 8.5 | ** |
| 9 | **** |
| 10 | * |
| 11 | *** |
| 12 | ** |
| 12.5 | ** |
| 13 | * |
| 14 | * |
| 15 | ** |
| 16 | * |
| 16.5 | * |
| 17 | * |
| 21.5 | * |
| 25 | * |
| 27 | * |
n = 7
| 0 | * |
| 1 | ** |
| 2 | ** |
| 4 | ** |
| Test name | n | r | p value |
|---|---|---|---|
| Chimera Test (Bill Bultas) | 5 | 0.96 | 0.05 |
| Tests by Kevin Langdon (aggregate) | 4 | 0.95 | 0.10 |
| The Nemesis Test | 5 | 0.91 | 0.07 |
| Numbers | 28 | 0.87 | 0.000006 |
| Ultra Test (Ronald K. Hoeflin) | 4 | 0.86 | 0.13 |
| Qoymans Multiple-Choice #1 | 4 | 0.85 | 0.14 |
| Isis Test | 4 | 0.83 | 0.15 |
| Analogies #1 | 4 | 0.83 | 0.15 |
| Analogies subtest of Long Test For Genius (Netherlandic) | 7 | 0.77 | 0.06 |
| Titan Test (Ronald K. Hoeflin) | 10 | 0.77 | 0.02 |
| The Final Test | 22 | 0.75 | 0.0006 |
| Space, Time, and Hyperspace | 24 | 0.74 | 0.0004 |
| Long Test For Genius (Netherlandic) | 7 | 0.73 | 0.07 |
| Long Test For Genius | 12 | 0.73 | 0.02 |
| Genius Association Test | 6 | 0.71 | 0.12 |
| Miller Analogies Test (before 2001, raw) | 4 | 0.68 | 0.25 |
| The Test To End All Tests | 7 | 0.58 | 0.16 |
| Mega Test (Ronald K. Hoeflin) | 15 | 0.56 | 0.04 |
| Cartoons of Shock | 5 | 0.52 | 0.30 |
| Chimera High Ability Riddle Test (Bill Bultas) | 10 | 0.51 | 0.12 |
| Analogies of Long Test For Genius | 14 | 0.47 | 0.09 |
| Miscellaneous tests | 28 | 0.46 | 0.02 |
| Hoeflin Power Test (Ronald K. Hoeflin) | 4 | 0.46 | 0.42 |
| Graduate Record Examination (prior to October 2001) | 6 | 0.41 | 0.36 |
| Association subtest of Long Test For Genius | 14 | 0.39 | 0.16 |
| Cooijmans Intelligence Test - Form 1 | 10 | 0.37 | 0.27 |
| W-87 (International Society for Philosophical Enquiry) | 5 | 0.27 | 0.60 |
| Bonsai Test | 5 | 0.27 | 0.60 |
| Spatial section of Test For Genius - Revision 2004 | 4 | 0.19 | 0.74 |
| Verbal section of Test For Genius - Revision 2004 | 4 | 0.12 | 0.84 |
| Raven's Advanced Progressive Matrices (I.Q.) | 8 | 0.07 | 0.84 |
| Association subtest of Long Test For Genius (Netherlandic) | 7 | -0.01 | 0.97 |
| Cattell Culture Fair | 8 | -0.11 | 0.76 |
| Culture Fair Numerical Spatial Examination - Final version (Etienne Forsström) | 4 | -0.46 | 0.42 |
| Qoymans Multiple-Choice #3 (batch scored by Paul Cooijmans) | 5 | -0.62 | 0.22 |
| Cattell Verbal (also known as Cattell B) | 4 | -0.97 | 0.10 |
Weighted mean of correlations: 0.527 (N = 317)
Estimated g factor loading: 0.73
Ranking in above table is based on the unrounded correlations. All available data is present in this table, no tests are left out except for those with less than 4 score pairs. All known pairs are used, including possible floor/ceiling scores or outliers.
These are estimated g factor loadings, but against homogeneous tests (containing only particular item types) as opposed to non-compound heterogeneous tests. Although tending to surprise the lay person, it is not uncommon for tests to have high loadings on item types they do not actually contain themselves. Such loadings reflect the empirical fact that most tests for mental abilities measure primarily g, regardless of their contents; that the major part of test score variance is caused by g, and only a minor part by factors germane to particular item types. It is of key importance to understand that this is a fact of nature, a natural phenomenon, and not something that was built into the tests by the test constructors.
| Type | n | g loading of Short Test For Genius on that type |
|---|---|---|
| Verbal | 112 | 0.67 |
| Numerical | 28 | 0.93 |
| Spatial | 44 | 0.64 |
| Heterogeneous | 82 | 0.73 |
N = 266
Compound tests have been left out of this table to avoid overlap.
Balanced g loading = 0.74
| Country | n | median score |
|---|---|---|
| Finland | 6 | 11.5 |
| Belgium | 3 | 11.0 |
| Netherlands | 7 | 7.0 |
| Sweden | 3 | 5.0 |
| United_States | 19 | 5.0 |
| United_Kingdom | 4 | 4.5 |
| France | 5 | 4.0 |
| Unknown | 17 | 4.0 |
Total number of countries: 16
For reasons of privacy, only countries with 3 or more candidates are included in this table. Ranking is based on the medians, and then alphabetic.
Notice: A correlation is generally considered significant if its p value is 0.05 or less.
| Personalia | n | r | p value |
|---|---|---|---|
| Observed behaviour | 12 | 0.79 | 0.009 |
| Gifted Adult's Inventory of Aspergerisms | 4 | 0.62 | 0.28 |
| Observed associative horizon | 14 | 0.54 | 0.05 |
| Educational level | 15 | 0.47 | 0.08 |
| Sex | 73 | 0.30 | 0.01 |
| Father's educational level | 11 | 0.29 | 0.37 |
| Mother's educational level | 11 | 0.08 | 0.78 |
| Year of birth | 63 | -0.01 | 0.92 |
| Disorders (own) | 13 | -0.10 | 0.74 |
| Disorders (parents and siblings) | 11 | -0.40 | 0.20 |
The goal of estimated g factor loadings for restricted ranges is to verify the hypothesis that g becomes less important, accounts for a smaller proportion of the variance, at higher I.Q. levels. The mere fact of restricting the range like this also depresses the g loading compared to computing it over the test's full range, so it would be normal for these values to be lower than the test's full-range g loading.
| Below 1st quartile (3.0) | 0.65 (N = 54) |
|---|---|
| Below median (5.0) | 0.47 (N = 110) |
| Above median (5.0) | 0.80 (N = 192) |
| Above 3rd quartile (9.0) | 0.69 (N = 96) |
The computation of Cronbach's alpha is problematic for this test because a small number of its numbered items were actually examples with the answers given, so that there is no variance for those items. The split-half reliabilities are probably underestimations of the true reliability for the same reason, and as a result, the error of measurement given below may be an overestimation of the actual error.
| Age class | n | Median score |
|---|---|---|
| 60 to 64 | 3 | 7.0 |
| 55 to 59 | 2 | 17.5 |
| 50 to 54 | 3 | 3.0 |
| 45 to 49 | 8 | 4.0 |
| 40 to 44 | 4 | 3.0 |
| 35 to 39 | 8 | 7.5 |
| 30 to 34 | 12 | 5.0 |
| 25 to 29 | 11 | 7.0 |
| 22 to 24 | 6 | 5.8 |
| 20 or 21 | 3 | 3.5 |
| 18 or 19 | 2 | 7.0 |
| 17 | 1 | 8.0 |
| 15 | 1 | 4.0 |
| Age unknown | 9 | 2.0 |
N = 73
| Year taken | n | Median score | protonorm |
|---|---|---|---|
| 1995 | 7 | 2.0 | 327 |
| 1996 | 17 | 4.0 | 368 |
| 1997 | 11 | 7.0 | 404 |
| 1998 | 18 | 6.0 | 388 |
| 1999 | 9 | 7.0 | 404 |
| 2000 | 2 | 6.3 | 396 |
| 2001 | 2 | 4.5 | 376 |
| 2002 | 2 | 11.5 | 465 |
| 2003 | 3 | 5.0 | 383 |
| 2004 | 2 | 6.3 | 396 |
N = 73
Item statistics are not published as that would help candidates. To detect bad items, answers and comments from candidates are studied, as well as, for each problem, the correlation with total score on the remaining problems (item-rest correlation) and the proportion of candidates getting it wrong (hardness of the item). Possible bad items are revised, replaced, or removed, possibly resulting in a revised version of the test.
| Verbal × Numerical | 0.65 | (p value: 0.00000004) |
| Verbal × Spatial | 0.65 | (p value: 0.00000004) |
| Numerical × Spatial | 0.62 | (p value: 0.0000001) |
Ideal values for correlations between sections are around .5, thus being a compromise between the test's ability to yield a "profile" and its ability to provide an indication of general intelligence. With a too high correlation (like .8 or higher) the sections measure basically the same so there is almost no profile information in them, with a too low correlation (like .2 or lower) the sections are so different that there is little point in combining them into a measure of general intelligence.
| Verbal | 0.81 |
| Numerical | 0.80 |
| Spatial | 0.80 |
Proportion of total score variance accounted for by this factor: 0.64
These first factor loadings represent the subtests' correlations with that which the test is primarily measuring, and are a somewhat better indicator than the subtests' correlations with total score, which is why the latter are not reported any more. The remainder of the total score variance may lie in other factors (group factors), specificity, and error.
Prop. = proportion of candidates outscored in this section. In parentheses the proportion outscored for any possible scores higher than the present score but lower than the next-higher score in the table.
| Score | Prop. | # scores (* = 1 score) |
|---|---|---|
| 0 | 0.048 (0.096) | ******* |
| 1 | 0.212 (0.329) | ***************** |
| 2 | 0.418 (0.507) | ************* |
| 3 | 0.568 (0.630) | ********* |
| 3.5 | 0.644 (0.658) | ** |
| 4 | 0.712 (0.767) | ******** |
| 5 | 0.808 (0.849) | ****** |
| 6 | 0.870 (0.890) | *** |
| 6.5 | 0.897 (0.904) | * |
| 7 | 0.918 (0.932) | ** |
| 7.5 | 0.938 (0.945) | * |
| 8 | 0.959 (0.973) | ** |
| 10 | 0.979 (0.986) | * |
| 13 | 0.993 (1.000) | * |
| Score | Prop. | # scores (* = 1 score) |
|---|---|---|
| 0 | 0.164 (0.329) | ************************ |
| 0.5 | 0.336 (0.342) | * |
| 1 | 0.438 (0.534) | ************** |
| 2 | 0.582 (0.630) | ******* |
| 2.5 | 0.644 (0.658) | ** |
| 3 | 0.678 (0.699) | *** |
| 3.5 | 0.705 (0.712) | * |
| 4 | 0.760 (0.808) | ******* |
| 5 | 0.822 (0.836) | ** |
| 6 | 0.890 (0.945) | ******** |
| 8 | 0.966 (0.986) | *** |
| 9 | 0.993 (1.000) | * |
| Score | Prop. | # scores (* = 1 score) |
|---|---|---|
| 0 | 0.068 (0.137) | ********** |
| 0.5 | 0.144 (0.151) | * |
| 1 | 0.315 (0.479) | ************************ |
| 1.5 | 0.493 (0.507) | ** |
| 2 | 0.651 (0.795) | ********************* |
| 2.5 | 0.801 (0.808) | * |
| 3 | 0.856 (0.904) | ******* |
| 3.5 | 0.925 (0.945) | *** |
| 4 | 0.959 (0.973) | ** |
| 5.5 | 0.979 (0.986) | * |
| 8 | 0.993 (1.000) | * |