As scientists' means of communication and information behaviors have evolved over the past several decades due to computing and networking developments, the field of library and information science (LIS) has responded by surveying their changing practices. This study continues this type of research but switches focus away from scientists' use of journal articles, the dominant means by which research libraries have supported their investigations, and instead concentrates on scientists' practices related to data. The view of scientists as increasingly energetic generators, managers, and users of large and growing electronic datasets has lately been recognized and promoted at the societal level by funding agencies, academic societies, and large research centers. The growth in data production and storage related to science, technology, engineering, and mathematics (STEM) research has been attributed to the development and proliferation of enabling technologies and computer networks associated with cyberinfrastructure advancement, including sensors and sensor networks, high-throughput technologies and instrumentation, automated data acquisition, and computational modeling and simulation. (NSF, 2007) While the technical and societal forces encouraging scientists to produce this born-digital content continue unabated, attention to their resulting burden of managing this content for access and use has lagged behind. LIS-trained practitioners could contribute information management infrastructure to aid scientists but this infrastructure would have to be oriented to the variety of local, idiosyncratic, field-based approaches currently extant. (Borgman, 2007) A campus-wide census of data management practices at a Carnegie Classification Research II institution identifies what the variety of faculty data management practices is across the STEM departments. LIS researchers have studied scientists to understand their information behaviors in a changing scholarly publishing environment. For instance, Brown (1999) polled faculty in four disciplines, Astronomy, Chemistry, Mathematics, and Physics, at her home institution to gather comparative data about their use of and attitudes about electronic versus print journal articles in relation to the academic library. Tenopir and her colleagues (2005) have teased out differences in the rate and manner that scientists in particular disciplines consume online versus print journal articles to conduct their research via multiple, sampled surveys that target a large number of members of the field's dominant academic society, and then comparing the results to other similar studies. Zhang (2001) surveyed authors of articles published in eight scholarly journals covering the LIS field to determine how Internet-based electronic resources were considered during the preparation of their articles. To help get the most out of computer-based data resulting from the increasing amount of scientific investigation facilitated through cyberinfrastructure, funding agencies in both the U.S. and the U.K. are investing in projects that develop data preservation and management tools and skills. One such effort, the Science Data Literacy (SDL) project at the iSchool at Syracuse University, has been funded by the National Science Foundation to develop a course designed to prepare students with basic knowledge and skills in science data management. In a similar manner as the earlier surveys of scientist information behaviors, the SDL staff as an early part of the project prepared a survey vehicle and conducted a census of the relevant campus faculty. The project team took a pragmatic approach in crossing disciplinary boundaries in order to gather the variety of data management practices in STEM departments across the home institution. Departments were identified that reasonably fell within the rough boundaries provided by the STEM category amalgam; SDL staff erred on the side of including the most number of researchers likely to be accumulating and working with primary datasets. In consultation with iSchool faculty practiced at survey construction, SDL staff created a survey vehicle and pilot-tested it on members of the SDL project advisory board who matched the target population. Terminology had to be negotiated and definitions provided to orient concepts from information science to the scholars' research process. For example, we eliminated the word “metadata” from the questionnaire and added an inclusive definition of data from a National Science Board report (2005) in the introduction to the survey: “any information that can be stored in digital form, including text, numbers, images, video or movies, audio, software, algorithms, equations, animations, models, simulations, etc.” Demographic questions particular to the discipline-focused and hierarchical environment of the academic community were taken from the Higher Education Research Institute's faculty performance survey (2004). As per Janes' (1999) helpful advice on survey construction, this demographic data-gathering section was placed at the end. The iterative process of question phrasing and grouping led to a refined instrument designed to capture data management practices. The four-part, web-based survey featured Likert-scale agreement questions related to attitudes, practices, and experience with research data. Two provocative questions about data management in the respondent's discipline were designed not only to motivate participants to express their views, but also to help establish the presupposition of the questionnaire that the respondent was a data producer (Martin, 2006). This rather brute force technique was designed to concentrate the mind of the respondent on their data management practices, but may have had unknown effects on the target population. An initial branching question may have been more effective at weeding out members of the STEM departments who, due to a teaching, practical, or theoretical orientation, did not manage data. In the meat of the survey, branching questions also would have allowed researchers to record more than one management or preservation practice, but the instrument provided for this purpose the ability for the participant to mark multiple options per response where appropriate to simplify the survey design. Textboxes and open-ended questions in each section were meant to elicit a range of practices with data as well as overall reactions provoked by the 25 questions. Syracuse University, in combination with the symbiotically connected campus of the SUNY College of Environmental Science and Forestry, has over a thousand full-time faculty members; the SDL survey was administered to 362 faculty members from STEM departments at the two institutions. Respondents received a movie ticket coupon as consideration for their time. Possible participants were identified via information posted on school and department websites and then contacted via email solicitation, notification and up to two reminders. Faculty members were given the opportunity to opt out—many did citing current retirement status or lack of involvement in research. As it was a local census using a MySQL-based token response system, anonymity was not an option, so entries are being kept in confidence and analysis is ongoing. Of those who participated resulting in a 30.7 percent response rate, problems that might affect the census results include a lack of alignment of the respondent with data-producing aspects of STEM research pursuits, such as a solely theoretical orientation in physics and mathematics, or a social orientation in the case of geography and information science. Additionally, faculty who handle different types of data per research project, faculty organized in research groups, and faculty use of research assistants as day-to-day handlers of data are conditions that may have interfered with an accurate and complete perception of data management practices from the survey responses obtained. Table 1 indicates some of the variety encountered in data management practices throughout the STEM departments. A value closer to five indicates strong agreement that the researcher responding is a frequent data producer, used data prepared by another researcher, was aware that the digital data may be used by another researcher outside his or her own group, prepared metadata of some kind for the internally produced datasets, and found metadata entries helpful when obtaining data for use from external research groups. All entries could have been the respondent answering as a representative of his or her research group. Table 2 shows the relation in responses between those researchers who work with a certain size dataset and their perception of the effect of data management practices on their discipline's progress. Researchers who operate with larger datasets appear more confident about their discipline's data management practices. Table 3 relates actions routinely taken by researchers with their data and with their agreement on the negative impact that inadequate data preservation practices are having on their discipline. It appears that researchers involved in advanced data activity such as calculation and visualization may be slightly more sanquine about their discipline's preservation practices. Continuing analysis will show the variations in management and preservation practice, according to faculty affiliation and position. This will allow mapping of local institutional attitudes and behaviors regarding data to those encouraged by major research initiatives at the disciplinary and national level. It is anticipated that results from this survey will help the SDL project team, and the LIS field more generally, understand how to integrate SDL education with STEM departments actively working with science data, as well as develop educational materials at an appropriate level to assist meeting the need for personnel with the skills and interest in science data management and preservation to support effective community and disciplinary use of these digital resources. On a more basic level, it will help provide information on how science and technology researchers obtain and manage their data as part of knowledge production and science communication processes.
No takes yet. Share an insight, caveat, or question.
D’Ignazio et al. (2008) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: