GRADUATE medical education is one of the core missions of academic medical centers, wherein medical specialists are responsible for teaching and supervising their future colleagues. However, being a medical specialist is no longer a sufficient qualification or proxy for competence in medical education aimed at training residents. This is particularly true given the modernization requirements for competency-based teaching and training promoted by accreditation institutions in some countries (such as the Accreditation for Council for Graduate Medical Education in the United States).1These modernization efforts accelerate faculty development of clinician-educators needed to achieve and maintain the highest standard of postgraduate medical education. An effective faculty development track should include measuring medical teaching effectiveness. This requires valid and reliable instruments, as well as providing the findings in a clear and concise format to faculty. Several studies have found that systematic and constructive feedback can result in improved teaching.2There are few published and validated evaluation systems or even instruments aimed at supporting the graduate medical education qualities of clinical faculty. In anesthesiology, there are few published instruments and systems,3and the existing ones tend to focus on faculty evaluation by residents only without any self-evaluations by faculty. To ensure actual behavioral change, individuals must usually undergo a stepwise change process. Evaluation insights obtained from feedback should be followed by creating positive intentions to change, trying out new behaviors and integrating them into practice. Supporting this change process has been shown to be effective.2,4To support the specialty-specific evaluation of teaching qualities of anesthesiology faculty in an academic medical center, we developed the System for Evaluation of Teaching Qualities (SETQ) comprising (1) a Web-based self-evaluation by faculty, (2) a Web-based residents’ evaluation of faculty, (3) individualized faculty feedback, and (4) individualized faculty follow-up support. This paper has three main objectives: (1) to investigate the psychometric properties of the two instruments underlying the SETQ system, (2) to explore the relationship between residents’ evaluation and faculty self-evaluation, and (3) to gauge the feasibility of reliably using residents’ evaluation of faculty by estimating the number of such evaluations needed per faculty. We also place these objectives in context by describing SETQ.SETQ was initially developed in the anesthesiology department of a large academic medical center that has over 7,000 staff (including about 500 faculty and 400 residents) in the Netherlands. It was later expanded to include specialty-specific modules for internal medicine, surgery, and obstetrics and gynecology. At the time of writing, most of the remaining specialties have signed up for SETQ, resulting in more than 90% faculty coverage in 2009. SETQ is receiving nationwide attention.Figure 1provides an overview of the SETQ system for evaluating teaching qualities of anesthesiology faculty. It was conceived as a three-stage, individualized measurement and improvement system. The first stage involved measurements using (1) a Web-based self-evaluation instrument filled in by faculty and (2) another Web-based instrument for evaluation of faculty by residents. The second stage involved individualized faculty feedback in which each participating faculty received detailed reports of the outcomes of the residents’ evaluations and, if available, self evaluations, also graphed within the context of the averaged outcomes of their colleagues. The third stage involved individualized faculty follow-up with the aim of discussing the results and finding avenues for improvement, if needed, with each individual faculty and head of department. The research team also provided anonymous overall feedback averaged over all participants to the entire departmental faculty and residents in medical teaching seminars.Data collection took place in the month of September 2008, when 33 residents who had been in training for at least 6 months and 39 faculty in the anesthesiology department were invited via e-mail to participate in the evaluations. The invitation assured the formative purpose and use of the evaluations. Participation of faculty and residents remained confidential. Residents’ evaluations were anonymous; only the number of residents’ evaluations was reported back to the individual faculty members. Each faculty was invited to share and discuss their feedback results with the head of department, but this was not mandatory. The two evaluation instruments were made available electronically via a dedicated password-protected SETQ Web portal. Residents chose who to evaluate and could evaluate many faculty. Each faculty could only self-evaluate. Automatic e-mail reminders and the head of department at clinical meetings encouraged both faculty and residents to participate in the evaluation.Both the self-evaluation and the residents’ evaluation instruments were based on the well-known 26-item Stanford Faculty Development Program (SFDP26) instrument, which was developed in the United States.5–9It is based on educational and psychological theories of learning and empirical observations of clinical teaching.In an earlier smaller study, we developed and pilot-tested the SFDP26 instrument for evaluation of anesthesiology faculty in an academic medical center outside the United States. A taskforce of anesthesiology faculty and residents drafted a questionnaire by translating the SFDP26 questionnaire and discussing its completeness, feasibility, and validity for a Dutch residency program. Consensus was reached. After its discussion in separate meetings of anesthesiology residents and faculty, the questionnaire was further edited, tested, and evaluated. We tentatively concluded that the adapted SFDP26 instrument completed by residents could yield reliable and valid evaluation of anesthesiology faculty in an academic medical center outside the United States.10Both the self-evaluation and residents’ evaluation instruments shared 24 core items spanning 5 domains of teaching quality, namely learning climate (8 items), professional attitude towards residents (4 items), communication of goals (4 items), evaluation (4 items), and feedback (4 items). Each of the 24 items had a 5-point Likert-type response scale: strongly disagree, disagree, neutral, agree, strongly agree. Each instrument concluded with two global rating items measuring “faculty being seen as a role model” and “faculty’s overall teaching qualities,” respectively. The global rating “faculty being seen as a role model” had the same response scale as the core items. For the global rating “faculty’s overall teaching qualities,” the 5-point Likert-type response was 1 = bad, 2 = fair, 3 = average, 4 = good, and 5 = excellent. In addition, the resident instrument had two open questions for narrative feedback on faculty, listing the strong teaching qualities of individual faculty and formulating concrete suggestions for improvement. We also collected data on residents’ sex and year of training. For faculty, we collected data on age, sex, number of years in practice since registration as an anesthesiologist, actual time spent on teaching residents, and previous participation in a training program for clinician-educators.We carried out four main types of analysis. First, we estimated the descriptive statistics (means, proportions) for the response sample to understand the basic characteristics of the participating residents and faculty.Second, to address the first objective of this study, that is, the psychometric properties of the SETQ instruments for both residents and faculty, we conducted exploratory factor, reliability coefficient, item-total scale correlation, interscale correlation, and scale versus global ratings correlation analyses.11,12For item reduction or multifactorial structuring of the instruments, we conducted exploratory factor analysis by using the principal components technique with oblique rotation to explore the factor or scale structure of both instruments separately. On the basis of the foregoing results, we calculated the internal consistency reliability coefficient or Cronbach’s α for each scale.13A Cronbach’s α of at least 0.70 was considered satisfactory.14,15In addition, we used the residents’ instrument to estimate faculty-level reliability coefficients of the intraclass correlation type based on the variance components for each scale.11Although no test-retest reliability was conducted in the current study, it was expected that any finding of high levels of interrater reliability would suggest that the intraobserver reliability, hence test-retest reliability, can only be higher.11Item-total scale correlations, corrected for item overlap, were used to check for the homogeneity of the scales based on averaging items that loaded strongly on the scales.11Furthermore, interscale correlations for residents and faculty separately were used to check for the interpretability of the constructed scales as distinct domains of a related overall construct. An interscale correlation of less than 0.70 was seen as satisfactory and gave credibility to the multidimensional factor or scale structure of the instruments.12,16To explore the construct validity of the instruments,11the scales were finally correlated with the two global ratings, “faculty being seen as a role model” and “faculty’s overall teaching qualities.” This approach was an imperfect, opportunistic construct validation testing using global rating items embedded within the instruments for construct hypothesis testing. We emphasize that this construct validation approach was not aimed at providing a final answer but at yielding initial results in an ongoing and cumulative exercise to be improved upon in subsequent research, as is increasingly acknowledged in the modern psychometrics literature.11Conceivably, an endless number of related hypotheses could be coined and tested for parts of the instruments over time. We hypothesized or made the informed assumption that faculty who score high on the items/scales should score highly on being seen as a role model and on the singular measure of their overall teaching qualities.17In line with the literature, we expected appropriate correlations between the scales and global ratings to fall within the range of 0.40 to 0.80.11Third, to investigate our second objective of exploring the relationship between residents’ assessments and faculty self-evaluations, we estimated the mean and SEM of each scale and their related items. Kendall’s rank order correlation coefficient τ18was used to gauge the correlations between the faculty’s teaching qualities rankings based on residents’ assessments versus those based on faculty self-evaluations. There is no generally accepted cut-point for high rank order correlation; for this study, therefore, the higher the correlation, the better.Fourth, this study’s final objective of investigating the feasibility of reliably using residents’ evaluation was analyzed by estimating the number of per-resident evaluations of faculty. This estimation involved solving the equation for the aforementioned reliability coefficient of the intraclass correlation type (using variance components) to determine the number of resident evaluations needed per faculty at any predefined reliability level.11,19,20We further triangulated the estimates of the number needed obtained above from solving the variance partitioning of the cross-classified multilevel model equation as follows. It was assumed that for each scale or instrument, the ratio of the sample size (N ) to the reliability coefficient (R ) would be approximately constant across combinations of sample size and associated reliability coefficients.11,21Therefore, the number of residents’ evaluations needed (N new) divided by the needed reliability coefficient (R new) would be equal to the observed number of residents’ evaluations per faculty (N old) divided by the observed reliability coefficient R old. We already knew N oldand R oldand could assume different target values for R new; therefore, we easily estimated N newfrom the assumed equality N new/R new=N old/R old. We repeated the calculations for reliability coefficients (R new) of 0.60, 0.70, 0.80, and 0.90. Reassuringly, both the first but complex and second but simple methods gave similar results within an error margin of no more than ±1. The results of the first method are reported here.Statistical significance was set at P < 0.05 (two-tailed). All analyses were conducted by using the general purpose statistical software SPSS version 16.0.2 (SPSS Inc., Chicago, IL) and Microsoft Office Excel 2003 SP3 (Microsoft Corporation, Redmond, WA).There were 30 residents and 36 anesthesiology faculty who participated in the study, yielding response rates of 91% and 92%, respectively (table 1). Two-thirds of residents and one-third of faculty participants were female. Residents from all but the last year of training were represented. Residents completed a total of 611 evaluations. There were about 20 evaluations per resident and nearly 16 evaluations per faculty member. Faculty reported being registered anesthesiologists for a mean of 12.7 yr. About 19% of faculty reported having enjoyed a formal training for clinician-educators. The actual time spent on teaching varied substantially among faculty members. Table 1gives an overview of participant characteristics.Explorative factor analysis yielded five teaching domains or scales for both instruments: learning climate, professional attitude towards residents, communication of goals, evaluation of residents, and feedback (table 2). Cronbach’s α for the internal consistency reliability was high for the residents’ instrument ranging from 0.89 for the scale “professional attitude towards residents” to 0.94 for both “communication of goals” and “evaluation of residents.” Cronbach’s α was lower for the faculty self-evaluation instrument, ranging from 0.57 for “learning climate” to 0.86 for “communication of goals.” As a result, all scales except “learning climate” in the faculty instrument achieved reliability coefficients above 0.70. Furthermore, the estimates of the faculty (group) level reliability of the residents’ instrument ranged from 0.86 (for “professional attitude towards residents”) to 0.93 (for the “evaluation of residents” domain).The item-total scale correlations were high for most items within their scales and in many cases higher for the residents’ instrument than for the faculty instrument (table 2). Nonetheless, three items (Q03, Q04, and Q08) on the faculty instrument displayed low item-total correlations. As shown in table 3, the interscale correlations for the residents’ instrument ranged from 0.19 (between “professional attitude towards residents” and “evaluation of residents”) to 0.66, P < 0.01 (between “communication of goals” and “evaluation of residents”). The faculty instrument displayed similar results, from 0.04 (between “professional attitude towards residents” and “evaluation of residents”) to 0.63 (between “learning climate” and “evaluation of residents,”P < 0.01). For both instruments, all interscale correlations were less than the 0.70 threshold mentioned in the methods section above.Table 4displays the bivariate correlations of each of the five scales with the two global ratings. For the residents’ instrument, all scales were significantly and positively correlated with the global ratings. The feedback scale had the highest correlations with global ratings of “faculty being seen as a role model” (0.60, P < 0.001) and “faculty’s overall teaching qualities” (0.68, P < 0.001). Contrastingly, the scale “professional attitude towards residents” had the lowest correlations with the global ratings (0.43 and 0.37, respectively, P < 0.001). For the faculty self-evaluation instrument, the “learning climate” scale had the highest correlation (0.66, P < 0.001) with the global rating “faculty being seen as a role model.” However, the “evaluation of residents” scale had the highest correlation (0.59, P < 0.001) with the global rating “faculty’s overall teaching qualities.” Overall, these correlations between the scales and global ratings tended to fall within the expected moderate range of 0.40 to 0.80, according to the literature.11Table 5shows that residents’ assessments of the teaching qualities of their anesthesiology faculty were positive. On scale of 5, the means of the residents’ evaluation scale scores for their faculty ranged from 3.41 for “communication of goals” to 4.15 for “professional attitude towards residents.” The faculty evaluated themselves highly, with their mean scale scores ranging from 3.22 for “communication of goals” to 4.13 for “professional attitude towards residents.”Looking at the mean scores across the five scales, there was no clear pattern of whether faculty consistently scored themselves higher than the residents scored them. Yet, three of the five scales (“learning climate,”“professional attitude towards residents,” and “evaluation of residents”) showed low to moderate correlations between the rankings produced by the residents’versus faculty self scores. The individual items displayed similar results. The residents scored the faculty higher on the two global ratings than the faculty did themselves. There were no rank correlations between residents’versus faculty self scores on the global ratings (table 5).For reliable feedback to faculty by using residents’ evaluations, the analysis showed that assuming a reliability coefficient of 0.70 for the entire instrument, at least four completed assessments per faculty would be required (table 6). Applying a stricter reliability coefficient of 0.80 would require as many as seven residents evaluating each faculty.This study demonstrates that the two instruments underlying SETQ seem reliable and valid for the evaluation of the teaching qualities of faculty in an academic medical center. Although there were no large differences in the mean scores from the residents’ and self-evaluation of the faculty, the rankings produced by those scores were only lowly to moderately correlated if at all. Finally, we that the of residents’ assessments per faculty this 4 to in a anesthesiology department such as the one in which this study was discussing the and of these a few study should be First, the number of faculty to the number of the faculty instrument could have to the lower reliability coefficient of 0.57 and factor observed for items Q04, and on the “learning climate” However, the lower factor could be expected for our sample to the items in the scale for psychometric observations could also be of faculty’s different of teaching to residents. Although these observations be a the psychometric of the items Q04, and should be in future research, given that items and were on a separate scale in the the of this study did not support of or or test-retest However, the high levels of interrater reliability found suggest that the intraobserver reliability can only be we observed that the three faculty who did not out their self-evaluation were scored lower than on all domains available from on This questions about the of on the of the findings for the faculty if residents to faculty who and the of faculty participation could be estimated it be to use the resulting to the faculty scores if the aim of measurement were to faculty the findings not be to residents and faculty in specialties or institutions each residency program and has its and is being to the findings of our studies in different given the follow-up of faculty, this study any about the of SETQ on the of anesthesiology are increasingly study, one developed that could be adapted for the systematic evaluation and support of faculty involved in those residency findings strong empirical support for the reliability and validity of the results obtained from the two and instruments for faculty The reliability findings as by the high faculty-level reliability coefficients estimated in table that residents’ instruments can be used to measure and the teaching qualities of faculty. The results of the psychometric analysis that we could into five domains seen as of teaching by both residents and faculty. We that the correlations among “learning of of residents,” and could that some items within those scales some For future we further the of the related expected correlations of each of the five scales with the two global ratings an support for the five teaching domains as of the of clinical This is on the assumption and and role for if the five scales teaching qualities and the global rating on overall teaching did we could the scales to at least moderately with the global rating was the The correlations should be high (for than that would to of the entire instrument, that is, if it could be to one global findings of the validity of the SETQ of strong correlations between faculty and residents’ ratings is with research and systematic that that had a to findings support the validity of the instruments among residents and faculty. The rank order correlations are not as a measure of the validity of the instruments but as a measure of the of between residents faculty’s teaching and the faculty themselves evaluate their It is for both instruments to be valid and reliable results among residents and faculty, respectively, and yield low correlations between faculty and residents as was found and self-evaluation from residents’ evaluation of faculty. We have no to that self-evaluation must or even moderately evaluation of the same in a this In based on results in behavioral of and given strong a and we would well to has that the as by that this be related to strong a and of qualities not shared by Nonetheless, the faculty self-evaluation used in those who are are increasingly being from different and in is feedback from such as residents, faculty focus their learning on improvement results also suggest that between four and residents’ evaluations per faculty the standard reliability level of 0.70 seen in the be required to faculty teaching qualities seem for most anesthesiology residency This finding the feasibility of the SETQ instruments was tested, and for formative It was as a the of the instruments and the of reliability would use in a high a system for measuring and, the of faculty teaching is it is to residents’ in clinical faculty feedback The individualized feedback reports in SETQ and reports from faculty suggest that SETQ was of effective teaching and to improvement. The reports faculty to focus their development and insights not be sufficient to about behavioral SETQ was to include a formative follow-up with the program at individual development studies suggest that this actual assessments to actual SETQ instruments provided reliable and valid results that could be used in formative support of anesthesiology faculty. This study further than previous to include the of the faculty to evaluate and The instruments are available, not require number of residents’ evaluations per faculty, and yield faculty feedback seen as in individual development the in this study of all anesthesiology faculty and residents of the Medical The Netherlands. The for with data and analysis.
No takes yet. Share an insight, caveat, or question.
Lombarts et al. (2009) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: