Item description in translated tests: Cultural effects on factor structure of the Chinese Version of the Myers-Briggs Type Indicator-G
Main Article Content
Construct validity is a core issue in the design of a standard psychological instrument. There are many contrasting arguments about the factor structure of the Myers-Briggs Type Indicator (MBTI, Myers, 1962). A variety of methods and sample sizes used in prior studies has created these differences. Therefore, this study explored factor structure of the Chinese version MBTI, Form G based on two samples with roughly equal sample sizes, using the same method. Results supported construct validity of the Chinese version of the MBTI-G (Miao, Huangfu, Chia, & Ren, 2004), for there were four factors extracted based on the two samples respectively. However, factor analysis results raised questions about the nature of the Sensation-Intuition scale. This study provided evidence that, for a collectivist culture, the description of items related to the individual’s own preference and attitude or other people’s value and attitude would affect the results of the Chinese version MBTI-G in construct validity evaluation.
The Myers-Briggs Type Indicator (MBTI; Myers, 1962) is a self-report, forced-choice personality inventory that was based on Carl Jung’s theory of psychological types (Jung, 1921/1971). Every year, between 1.5 million and 2 million people in the United States take the MBTI, and it has become the most widely used personality assessment tool (Goby & Lewis, 2000; Zemke, 1992). A major reason for the popularity of the MBTI is its relevance in many quite diverse areas − education, career development, organizational behavior, group functioning and team development, psychotherapy with individuals and couples, and in multicultural settings (Quenk, 1999). Because of its position as the most widely used scale for “normal” adults in diverse areas, the MBTI has been translated into many different languages such as Norwegian, Italian, and Chinese (Miao, Huangfu, Chia, & Ren, 2000; Nordvik, 1994; Saggino & Kline, 1996). With the translation and modification of the MBTI in Mainland China, there is a palpable need to explore the factor structure of the Chinese version MBTI, and this was the purpose of our study.
Based on the theory of psychological types, the MBTI consists of four bipolar scales: The Extraversion-Introversion (EI) scale which measures how an individual distributes his/her energy − directing mainly toward the outer world of people and objects as opposed to the inner world of ideas and experiences; The Sensing-Intuition (SN) scale which measures how one prefers to gather information − focusing mainly on the five senses as opposed to insight; The Thinking-Feeling (TF) scale which measures how one is likely to make a decision − based on logical analysis as opposed to needs for affiliation and warmth; the
Judging-Perceiving (JP) scale which measures how one chooses to approach life − preferring order and rules as opposed to flexibility and spontaneity. The eight raw scores, that are the scores on the two opposing poles of the four scales, can sort individuals into 16 possible rational categories or types.
A theory-based test must demonstrate that it adequately reflects the theory it purports to represent. Thus, there has been a wealth of research conducted on the construct validity of the MBTI in recent decades. However, the number of factors extracted varied in these studies. For example, in the Comrey (1983) study, the TF scale was divided into two factors to create a fifth factor. In the Sipps, Alexander, and Friedt (1985) study, the first factor was a TF/SN combination scale. The sixth factor was a very weak scale that did not seem to have a distinguishing characteristic. In the Harasym, Leong, Juschka, Lucier, and Lorscheider (1996) study, the results showed that the MBTI consisted of three bipolar scales (EI, TF and SJ-NP). In Saggino and Kline’s 1995 study, an important difference from other studies concerned the appearance of the fifth factor which mainly consisted
of word-pair items.
There seem to be at least two main reasons for the diverse number of factors obtained. Firstly, the methodological differences in prior studies, which can create artificial results inseparable from the actual results (Tischler, 1994). Secondly, sample sizes have covered a rather broad range from 241 (Comrey, 1983) to 2143 (Tischler). In fact, most researchers believe that factor analysis should have at least a 5:1 ratio, and preferably a 10:1 ratio, of subjects to items (Tischler). Accordingly, a 10:1 ratio for 94 items would mean a sample of 940+ and a 5:1 ratio would mean a sample of 470+. However, some previous studies had fewer participants than required for a 10:1 ratio criteria, or a sample even smaller than the 5:1 ratio criterion required. Therefore, the small sample size in these studies might have resulted in the variability of factor analysis (Tischler).
Considering the effects on factor structure mentioned above, in this study we used the same method with two larger samples to compare the factor structure of the Chinese version MBTI.
Method
Participants
Sample A consisted of 1,765 juniors in military universities (1,455 males and 310 females) ranging in age from 17 to 26, with an average age of 20 (SD = 1.50). Sample B consisted of 2,290 students who applied for military university (1,936 males and 354 females) ranging in age between 16 to 20 years, with an average age of 18.3 (SD = 0.70). Both samples fall into the same age range of young adulthood (18-22).
Procedure
The Chinese version of the MBTI-G was administered to samples A and B, separately. Individuals in both groups were asked to complete the instrument in line with instructions given. It was explained to sample A that the results of the MBTI would have no impact on their grades or performance, because the aim of the psychological test was just to collect data for research. However, sample B were told that the results of the MBTI would impact on their military university enrolments, because the instrument was used as one of the selecting tools.
Analysis
Initially, assumptions for factor analysis were examined to ensure that the two particular data sets were suitable for diagonalization. Specifically, the anti-image covariance matrices of two data sets had off-diagonal values approaching zero (as are appropriate), the Kaiser-Meyer-Olkin tests of sampling adequacy values (0.851 and 0.861 respectively) were large (as anticipated), and Bartlett’s sphericity tests were found to be significant (p = .000). These findings indicated that the two data sets were factorable and that eventual findings could be
supported by the data.
Confirmatory factor analysis (CFA) is becoming increasingly popular as a test of the validity of instruments. However, CFA is set up for normally distributed continuous data, not for dichotomous data, and MBTI items are not designed to predict an individual’s position on any scale but simply to show which pole of a bipolar scale they are on (Tischler, 1994). For this reason, it was inappropriate to use CFA in this study. Therefore, principal axis factoring and direct oblimin rotation (a kind of oblique rotation) were used to conduct the exploratory factor analysis. Oblique rotation was selected as the best method of rotation because the factors have an oblique relation to one another by both theory and prior empirical analyses. All statistical analyses were conducted with the SPSS for Windows, Version 13.0.
Results
Determining the number of factors extracted is a fundamental issue in exploring factor analysis. A common practice is to terminate extraction when the eigenvalues for successive factors are less than 1, begin to level off, or begin explaining similar proportions of variance from total variance explained (Kaiser, 1960; Thompson & Borrello, 1986). According to the data set of sample A, there were 28 factors with eigenvalues greater than 1. However, a scree plot of eigenvalues indicated that meaningful factors were few. The eigenvalues for the first seven factors were: 5.86, 5.57, 4.00, 2.27, 1.93, 1.78, 1.70. Based on Cattell’s scree test and the eigenvalues, it could be seen that the trailing off (scree) began with the fifth factor, leaving only the first four factors as the “true” factors; they accounted for 18.80% of the variance. Given various results of prior studies, we also extracted three and five factors for comparison with the four-factor model. Similarly, according to the data set of sample B, there were 27 factors with eigenvalues greater than 1; the first four factors were extracted, which accounted for 17.82% of the variance.
The factor-loading structure matrices of the two samples are presented in Table 1. Two findings were evident in this table. Firstly, regardless of whether three, four or five factors were extracted, all EI items always loaded on one factor based on the two samples. Secondly, most SN items were split into two factors in the three- and five-factor model based on sample A, and they mainly loaded on one factor in the corresponding model while using sample B.
In order to establish the reason why the SN scale split into two different factors, and any other scales did not in some models, we carefully compared each scale’s items. The interesting finding was that items could be roughly categorized into two subtypes in terms of description of items in the SN scale. Description of SN items in one subtype is mainly related to an individual’s own preference and attitude. But, in the other subtype, it is mainly related to other people’s values and attitude. Nevertheless, most items are described consistently within one context in other scales. Table 2 presents the distribution of items with different statements in each scale.
Table 1. Comparing the Distribution of Items on Correct Factors and Incorrect Factors in Different Models
Note: EI = Extraversion-Introversion Scale; SN = Sensation-Intuition Scale; TF = Thinking-Feeling Scale; JP = Judging-Perceiving Scale.
an = 1,765. bn = 2,290.
Table 2. Distribution of Items with Different Statement in Each Scale
Note: EI = Extraversion-Introversion Scale; SN = Sensation-Intuition Scale; TF = Thinking-Feeling Scale; JP = Judging-Perceiving Scale.
Discussion
Interest in the MBTI has stimulated a number of validity studies. In this study, we explored the factor structure of the Chinese version MBTI. According to the factor retention criterion, four factors should be extracted based on either sample A or B, which was in accordance with Jung’s theory. The cumulative variance explained by these four factors may have been low in this study, but it was similar with prior studies. For example, according to the study by Osterlind, Miao, Sheng, and Chia (2004), the first four factors accounted for 13.78% and 19.13% of the variance in the 1997 and 1999 Chinese translations, respectively. For the original English version, 25.40% of cumulative variance was explained by the first four factors (Tischler, 1994). In the study of Saggino and Kline (1996), the first five factors accounted for 20.20% of the variance in the Italian version MBTI.
Comparing the two four-factor models based on samples A and B, there were a few items that failed to load on their hypothesized “correct” factors. These items, distributed in the SN, TF, and JP scales, may be evidence that, in the Chinese version of the MBTI-G, these three scales are not as pure as the original theory predicted. This result was in line with prior studies. Harvey, Murry, and Stamoulis (1995) indicated that although the analyses strongly supported a four-factor model with Form G, the necessity of 11 items that had additional secondary loadings must be considered (most of these 11 items were in the SN scale). Results from the Sipps et al. (1985) study apparently suggested that the SN and TF scales were not factorially pure. In the Italian version of the MBTI Form F, the results indicated that the MBTI SN and TF scales were not pure: they had absolute loadings greater than 0.30 in more than one factors (Saggino & Kline, 1996). Unlike the other three dimensions, the EI scale was totally consistent with Jung’s theory. Whether three, four or five factors were extracted, all EI items always loaded on one factor based on the two samples. This result supported the findings of previous researchers’.
The most interesting finding in this study was about the SN scale. When subjects were military academy cadets, most SN items could be classified into two categories. However, when subjects were students aiming to be cadets, the two categories were integrated to form the SN factor. We questioned the nature of the SN scale because of this phenomenon. Why did differences exist in the SN scale between the two samples with the same factor analysis method and sample size? According to Quenk (1999), there are many factors that might influence participants’ self-report on the MBTI. In our study, the distinct difference between the two samples was testing context. Specifically, participants in sample B were in a selecting context, for their results of the MBTI would influence their enrolment. Nevertheless, the MBTI results of participants in sample A were presented as solely for research purposes.
Why is it that the SN scale’s structure could vary so easily in different testing contexts? After analyzing SN items in the Chinese version of the MBTI-G we found that during the process of amendment of the original instrument, the researchers described SN items in two different ways. Specifically, statements of SN items could be related mainly to the individual’s own preferences and attitudes or to other people’s values and attitude. An example in the “personal” category is as follows:
In reading for pleasure, do you
(a) enjoy odd or original ways of saying things, or
(b) like writers to say exactly what they mean?
An example in the “people in general” category is: Would you rather be considered
(a) a practical person, or
(b) an ingenious person?
Considering the cultural emphasis specific to Chinese people, different structures of the SN scale based on the two independent samples may be a feasible way to explain the differences. China is well-known as a collectivist country (Oyserman, Coon, & Kemmelmeier, 2002). In collectivist cultures one’s identity is defined more in relation to others. In other words, people of China consider objects in their relationship with people and the environment. Therefore, self-concept in collectivist cultures is malleable (context-specific) rather than stable (enduring across situations; Myers, 2005).
To meet the expectations of the future study environment, participants in sample B had a perception that a certain kind of person was desired for admission into the military university. In most people’s minds, military university is a place with rigorous rules and regulations. What cadets can do and what they should do every day is highly prescriptive. Members of such a university should obey rules and principles totally; individuality is not allowed. In order to be accepted into this environment, participants in sample B would tend instinctively to meet a concept of an “ideal personality type” that they had formed in their mind. This is a good example of a social response mind set. Although some applicants had probably not known about the real environment of a military university, they unconsciously tended to answer all SN items with the response they assumed to be desirable, regardless of the items’ descriptions related to individual preference or the values and attitudes of people in general. As a result, most of the SN items loaded on one factor based on the data set of sample B. Conversely, the same tendency did not appear with participants in sample A for two reasons: 1) They were already in the university and there was no tension about being admitted; and 2) they were already familiar with their campus environment and cadets knew that individuality could be present on appropriate occasions, so they could answer the test objectively. As a result, the SN items with descriptions related to individual preferences and attitudes separated from other SN items based on the data set of sample A.
However, the conclusion that the SN scale could be divided into two parts conflicted with a prior study in which the SN scale was divided into three subtypes (ideal and reality; abstract and concrete; creativity and convention; Osterlind et al., 2004). This two vs. three subtype inconsistency could be explained by the fact that the three subtypes are based on the content of SN items while the two subtypes are based on descriptions of SN items.
In the present study, with the same factor analysis method and the same instrument, we explored the two independent samples with roughly equivalent sample sizes to compare their factor analysis results. Essentially, four factors were extracted based on each of the two samples. This result supported the construct validity of this Chinese version of the MBTI-G. However, results also raised questions as to the nature of the SN scale. Differences in results from our two samples could be attributed to different motivations during the period of test, with sample B having stronger social desirability. But our interpretation was that it is more likely to have been caused by the distinctive collectivist culture of the Chinese people. If this collectivist culture had such a dominant influence, this would remind all researchers that, before using standardized psychological instruments that were originally developed in other languages, the most important thing for psychologists to take into consideration is the influence of cultural differences. For the Chinese version of the MBTI, cross-cultural amendment should involve at least three aspects. Firstly, the meaning of every character: It is not adequate to have an “exact” literal translation, researchers must instead look for a functional equivalent in that particular culture. Secondly, the item type: The phrase item or the word-pair item figures prominently in how participants may interpret the item. Both these aspects were considered in a previous study on the Chinese MBTI (Osterlind et al., 2004). Thirdly, the description of the item: The present study provided evidence that, in the Chinese collectivist culture, description of items translated into Chinese have a great impact on the results in construct validity evaluation. For this reason, when a standardized psychological instrument is used that was originally developed in another culture and written in another language, items with descriptions which would induce skewed responses in cultures should be amended to improve the accuracy of the instrument in selection.
One of the limitations of this study was that the sampling was restricted in types of participants. And the findings may not necessarily be generalizable to other groups in society. Therefore, further analyses with different samples are desirable to verify the effect of item description on the SN scale.
Comrey, A. L. (1983). An evaluation of the Myers-Briggs Type Indicator. Academic Psychology Bulletin, 5, 115-129.
Goby, V. P., & Lewis, J. H. (2000). Using experiential learning theory and the Myers-Briggs Type Indicator in teaching business communication. Business Communication Quarterly, 63(3), 39-48.
Harasym, P. H., Leong, E. J., Juschka, B. B., Lucier, G. E., & Lorscheider, F. L. (1996). Relationship between Myers-Briggs Type Indicator and Gregorc Style Delineator. Perceptual and Motor Skills, 82, 1203-1210.
Harvey, R. J., Murry, W. D., & Stamoulis, D. T. (1995). Unresolved issues in the dimensionality of the Myers-Briggs Type Indicator. Educational and Psychological Measurement, 55(4), 535-544.
Jung, C. G. (1921/1971). Psychological types. Collected works (Vol. 6, R. F. C. Hull, Trans.). Princeton, NJ: Princeton University Press.
Kaiser, H. F. (1960). The application of electronic computers to factor analysis. Educational and Psychological Measurement, 20, 141-151.
Miao, D. M., Huangfu, E., Chia, R. C., & Ren, J. J. (2000). The validity analysis of the Chinese version MBTI. Acta Psychologica Sinica, 32(3), 324-331.
Myers, D. G. (2005). Social psychology. New York: McGraw-Hill.
Myers, I. B. (1962). Manual: The Myers-Briggs Type Indicator. Palo Alto, CA: Consulting Psychologists Press.
Nordvik, H. (1994). Two Norwegian versions of the MBTI, Form G: Scoring and internal consistency. Journal of Psychological Type, 29, 24-31.
Osterlind, S. J., Miao, D. M., Sheng, Y. Y., & Chia, R. C. (2004). Adapting item format for cultural effects in translated tests: Cultural effects on construct validity of the Chinese versions of the MBTI. International Journal of Testing, 4(1), 61-73.
Oyserman, D., Coon, H. M., & Kemmelmeier, M. (2002). Rethinking individualism and collectivism: Evaluation of theoretical assumptions and meta-analyses. Psychological Bulletin, 128(1), 3-72.
Quenk, N. L. (1999). Essentials of Myers-Briggs Type Indicator Assessment. New York: John Wiley & Sons, Inc.
Saggino, A., & Kline, P. (1995). Item factor analysis of the Italian version of the Myers-Briggs Type Indicator. Personality and Individual Differences, 19, 243-249.
Saggino, A., & Kline, P. (1996). The location of the Myers-Briggs Type Indicator in personality factor space. Personality and Individual Differences, 21(4), 591-597.
Sipps, G. J., Alexander, R. A., & Friedt, L. (1985). Item analysis of the Myers-Briggs Type Indicator. Educational and Psychological Measurement, 45, 789-796.
Thompson, B., & Borrello, G. M. (1986). Construct validity of the Myers-Briggs Type Indicator. Educational and Psychological Measurement, 45, 745-752.
Tischler, L. (1994). The MBTI factor structure. Journal of Psychological Type, 31, 24-31.
Zemke, R. (1992). Second thoughts about the MBTI. Training, 29(4), 43-49.
Table 1. Comparing the Distribution of Items on Correct Factors and Incorrect Factors in Different Models
Note: EI = Extraversion-Introversion Scale; SN = Sensation-Intuition Scale; TF = Thinking-Feeling Scale; JP = Judging-Perceiving Scale.
an = 1,765. bn = 2,290.
Table 2. Distribution of Items with Different Statement in Each Scale
Note: EI = Extraversion-Introversion Scale; SN = Sensation-Intuition Scale; TF = Thinking-Feeling Scale; JP = Judging-Perceiving Scale.
The authors would like to thank Xu-feng Liu
Wei Xiao
Jing-jing Gong
Sheng-jun Wu
Yun-feng Sun
Hai Yang
and Wei Wang of the Fourth Military Medical University of China for their invaluable help in data collection.
Appreciation is due to anonymous reviewers.