DataShare upgraded to v2.3 – The embargo enhancement release

Posted on December 7, 2016 by Pauline Ward

The latest upgrade of Edinburgh DataShare, from version 2.2 to 2.3, brings in several useability improvements.

Embargo expiry reminder
If you want to deposit your data in DataShare, but you want to impose a delay before your files become freely downloadable, you can apply an embargo to your submission – see our “Checklist for deposit” for a fuller explanation of the embargo feature. As of DataShare v2.3, if you apply an embargo to your deposit, DataShare will now send you an email reminder one week before the embargo is due to expire. This gives you time to make us aware if you need the embargo to be extended, or to send us the details of your paper if it has been published, so that we can add those to the metadata, to help users understand your data.
DOI added to the citation field immediately
When your DataShare deposit is approved by the curator, the system mints a new DOI for you. As of version 2.3, DataShare now immediately appends the URL containing that DOI into the â€œCitationâ€� field, which is visible at the top of the summary view page of your item. The â€œCitationâ€� field makes it easy for others to cite your data, because it provides them with text which they can copy and paste into any manuscript (or any other document where they want to cite the data). Previously you would have had to click on â€œShow full item recordâ€� to look for the DOI in the â€œPersistent identifierâ€� field, or wait for an overnight script to paste the DOI onto the end of the â€œCitationâ€� field.

Tombstone records
We now have the ability to leave a â€˜tombstoneâ€™ record in place for any DataShare item that is withdrawn. We only withdraw items in exceptional circumstances â€“ for example where there is a substantive error or omission in the data, such that we feel merely labelling the item as â€œSupersededâ€� is not sufficient. Now, when we tombstone an item, the files become unavailable indefinitely, but the metadata remain visible at the DOI and handle URLs. Whereas until now, every withdrawn item has become completely invisible, so that the original DOI and handle URLs produced a â€˜not foundâ€™ error.

Screenshot of a DataShare item's citation field with the DOI

Cortical parcellation citation – now with DOI!

Enjoy!

Pauline Ward

Research Data Service

Publishing Data Workflows (Guest Post by Angus Whyte)

Posted on March 8, 2016 by Pauline Ward

In the first week of March the 7^th Plenary session of the Research Data Alliance got underway in Tokyo. Plenary sessions are the fulcrum of RDA activity, when its many Working Groups and Interest Groups try to get as much leverage as they can out of the previous 6 months of voluntary activity, which is more usually coordinated through crackly conference calls.

The Digital Curation Centre (DCC) and others in Edinburgh contribute to a few of these groups, one being the Working Group (WG) on Publishing Data Workflows. Like all such groups it has a fixed time span and agreed deliverables. This WG completes its run at the Tokyo plenary, so thereâ€™s no better time to reflect on why DCC has been involved in it, how weâ€™ve worked with others in Edinburgh and what outcomes itâ€™s had.

DCC takes an active part in groups where we see a direct mutual benefit, for example by finding content for our guidance publications. In this case we have a How-to guide planned on â€˜workflows for data preservation and publicationâ€™. The Publishing Data Workflows WG has taken some initial steps towards a reference model for data publishing, so it has been a great opportunity to track the emerging consensus on best practice, not to mention examples we can use.

One of those examples was close to hand, and DataShareâ€™s workflow and checklist for deposit is identified in the report alongside workflows from other participating repositories and data centres. That report is now available on Zenodo. [1]

In our mini-case studies, the WG found no hard and fast boundaries between â€˜data publishingâ€™ and what any repository does when making data publicly accessible. Itâ€™s rather a question of how much additional linking and contextualisation is in place to increase data visibility, assure the data quality, and facilitate its reuse. Hereâ€™s the working definition we settled on in that report:

Research data publishing is the release of research data, associated metadata, accompanying documentation, and software code (in cases where the raw data have been processed or manipulated) for re-use and analysis in such a manner that they can be discovered on the Web and referred to in a unique and persistent way.

The â€˜key componentsâ€™ of data publishing are illustrated in this diagram produced by Claire C. Austin.

Data publishing components. Source: Claire C. Austin et al [1]

As the Figure implies, a variety of workflows are needed to build and join up the components. They include those â€˜upstreamâ€™ around the data collection and analysis, â€˜midstreamâ€™ workflows around data deposit, packaging and ingest to a repository, and â€˜downstreamâ€™ to link to other systems. These downstream links could be to third-party preservation systems, publisher platforms, metadata harvesting and citation tracking systems.

The WG recently began some follow-up work to our report that looks â€˜upstreamâ€™ to consider how the intent to publish data is changing research workflows. Links to third-party systems can also be relevant in these upstream workflows. It has long been an ambition of RDM to capture as much as possible of the metadata and context, as early and as easily as possible. That has been referred to variously as â€˜sheer curationâ€™ [2], and â€˜publication at source [3]). So we gathered further examples, aiming to illustrate some of the ways that repositories are connecting with these upstream workflows.

Electronic lab notebooks (ELN) can offer one route towards fly-on-the-wall recording of the research process, so the collaboration between Research Space and University of Edinburgh is very relevant to the WG. As noted previously on these pages [4] ,[5], the RSpace ELN has been integrated with DataShare so researchers can deposit directly into it. So we appreciated the contribution Rory Macneil (Research Space) and Pauline Ward (UoE Data Library) made to describe that workflow, one of around half a dozen gathered at the end of the year.

The examples the WG collected each show how one or more of the recommendations in our report can be implemented. There are 5 of these short and to the point recommendations:

Start small, building modular, open source and shareable components
Implement core components of the reference model according to the needs of the stakeholder
Follow standards that facilitate interoperability and permit extensions
Facilitate data citation, e.g. through use of digital object PIDs, data/article linkages, researcher PIDs
Document roles, workflows and services

The RSpace-DataShare integration example illustrates how institutions can follow these recommendations by collaborating with partners. RSpace is not open source, but the collaboration does use open standards that facilitate interoperability, namely METS and SWORD, to package up lab books and deposit them for open data sharing. DataShare facilitates data citation, and the workflows for depositing from RSpace are documented, based on DataShareâ€™s existing checklist for depositors. The workflow integrating RSpace with DataShare is shown below:

RSpace-DataShare Workflows

For me one of the most interesting things about this example was learning about the delegation of trust to research groups that can result. If the DataShare curation team can identify an expert user who is planning a large number of data deposits over a period of time, and train them to apply DataShareâ€™s curation standards themselves they would be given administrative rights over the relevant Collection in the database, and the curation step would be entrusted to them for the relevant Collection.

As more researchers take up the challenges of data sharing and reuse, institutional data repositories will need to make depositing as straightforward as they can. Delegating responsibilities and the tools to fulfil them has to be the way to go.

[1] Austin, C et al.. (2015). Key components of data publishing: Using current best practices to develop a reference model for data publishing. Available at: http://dx.doi.org/10.5281/zenodo.34542

[2] â€˜Sheer Curationâ€™ Wikipedia entry. Available at: https://en.wikipedia.org/wiki/Digital_curation#.22Sheer_curation.22

[3] Frey, J. et al (2015) Collection, Curation, Citation at Source: Publication@Source 10 Years On. International Journal of Digital Curation. 2015, Vol. 10, No. 2, pp. 1-11

http://doi:10.2218/ijdc.v10i2.377

[4] Macneil, R. (2014) Using an Electronic Lab Notebook to Deposit Data http://datablog.is.ed.ac.uk/2014/04/15/using-an-electronic-lab-notebook-to-deposit-data/

[5] Macdonald, S. and Macneil, R. Service Integration to Enhance Research Data Management: RSpace Electronic Laboratory Notebook Case Study International Journal of Digital Curation 2015, Vol. 10, No. 1, pp. 163-172. http://doi:10.2218/ijdc.v10i1.354

Angus Whyte is a Senior Institutional Support Officer at the Digital Curation Centre.

Fostering open science in social science

Posted on June 24, 2015 by Rocio von Jungenfeld

On 10th of June, the Data Library team ran two workshops in association with the EU Horizon 2020 project, FOSTER (Facilitate Open Science Training for European Research), and the Scottish Graduate School of Social Science.

The aim of the morning workshop, â€œGood practice in data management & data sharing with social research,â€� was to provide new entrants into the Scottish Graduate School of Social Science with a grounding in research data management using our online interactive training resource MANTRA, which covers good practice in data management and issues associated with data sharing.

The morning started with a brief presentation by Robin Rice on ‘open science’ and its meaning for the social sciences. Pauline Ward then demonstrated the importance of data management plans to ensure work is safeguarded and that data sharing is made possible. I introduced MANTRA briefly, and then Laine Ruus assigned different MANTRA units to participants and asked them to briefly go through the units and extract one or two key messages and report back to the rest of the group. After the coffee break we had another presentation on ethics, informed consent and the barriers for sharing, and we finished the morning session with a ‘Do’s and Dont’s exercise where we asked participants to write in post-it notes the things they remembered, the things they were taking with them from the workshop: green for things they should DO, and pink for those they should NOT. Here are some of the points the learners posted:

DO
– consider your usernames & passwords
– read the Data Protection Act
– check funder/institution regulations/policies
– obtain informed consent
– design a clear consent form
– give participants info about the research
– inform participants of how we will manage data
– confidentiality
– label your data with enough info to retrieve it in future
– develop a data management plan
– follow the certain policies when you re-use dataset[s] created by others
– have a clear data storage plan
– think about how & how long you will store your data
– store data in at least 3 places, in at least 2 separate locations
– backup!
– consider how/where you back up your data
– delete or archive old versions
– data preservation
– keep your data safe and secure with the help of facilities of fund bodies or university
– think about sharing
– consider sharing at all stages. Think about who will use my data next
– share data (responsibly)

DONâ€™T
– unclear informed consent
– a sense of forcing participants to be part of research
– do not store sensitive information unless necessary
– don’t staple consent forms to de-identified data records/store them together
– take information security for granted
– assume all software will be able to handle your data
– don’t assume you will remember stuff. Document your data
– assume people understand
– disclose participants’ identity
– leave computer on
– share confidential data
– leave your laptop on the bus!
– leave your laptop on the train!
– leave your files on a train!
– don’t forget it is not just my data, it is public data
– forget to future proof

Our message was that open science will thrive when researchers:

organise and version their data files effectively,
provide comprehensive and sufficient documentation for others to understand and replicate results and thus cite the source properly
know how to store and transport your data safely and securely (ensuring backup and encryption)
understand legal and ethical requirements for managing data about human subjects
Recognise the importance of good research data management practice in your own context

The afternoon workshop on â€œOvercoming obstacles to sharing data about human subjectsâ€� built on one of the main themes introduced in the morning, with a large overlap of attendees. The ethical and regulatory issues in this area can appear daunting. However, data created from research with human subjects are valuable, and therefore are worth sharing for all the same reasons as other research data (impact, transparency, validation etc). So it was heartening to find ourselves working with a group of mostly new PhD students, keen to find ways to anonymise, aggregate, or otherwise transform their data appropriately to allow sharing.

Robin Rice introduced the Data Protection Act, as it relates to research with human subjects, and ethical considerations. Naturally, we directed our participants to MANTRA, which has detailed information on the ethical and practical issues, with specific modules on â€œData protection, rights & accessâ€� and â€œSharing, preservation & licensingâ€�. Of course not all data are suitable for sharing, and there are risks to be considered.

In many cases, data can be anonymised effectively, to allow the data to be shared. Richard Welpton from the UK Data Archive shared practical information on anonymisation approaches and tools for ‘statistical disclosure control’, recommending sdcMicroGUI (a graphical interface for carrying out anonymisation techniques, which is an R package, but should require no knowledge of the R language).

DrNiamhMoore Finally Dr Niamh Moore from University of Edinburgh shared her experiences of sharing qualitative data. She spoke about the need to respect the wishes of subjects, her research gathering oral history, and the enthusiasm of many of her human subjects to be named in her research outputs, in a sense to own their own story, their own words.

Links:

Rocio von Jungenfeld & Pauline Ward
EDINA and Data Library

Open up! On the scientific and public benefits of data sharing

Posted on January 22, 2015 by Robin Rice

Research published a year ago in the journal Current Biology found that 80 percent of original scientific data obtained through publicly-funded research is lost within two decades of publication. The study, based on 516 random journal articles which purported to make associated data available, found the odds of finding the original data for these papers fell by 17 percent every year after publication, and concluded that â€œPolicies mandating data archiving at publication are clearly neededâ€� (http://dx.doi.org/10.1016/j.cub.2013.11.014).

In this post Iâ€™ll touch on three different initiatives aimed at strengthening policies requiring publicly funded data â€“ whether produced by government or academics â€“ to be made open. First, a report published last month by the Research Data Alliance Europe, â€œThe Data Harvest: How sharing research data can yield knowledge, jobs and growth.â€�Â Second, a report by an EU-funded research project called RECODE on â€œPolicy Recommendations for Open Access to Research Dataâ€�, released last week at their conference in Athens.Â Third, the upcoming publication of Scotlandâ€™s Open Data Strategy, pre-released to attendees of an Open Data and PSI Directive Awareness Raising Workshop Monday in Edinburgh.

Experienced so close together in time (having read the data harvest report on the plane back from Athens in between the two meetings), these discrete recommendations, policies and reports are making me just about believe that 2015 will lead not only to a new world of interactions in which much more research becomes a collaborative and integrative endeavour, playing out the idea of â€˜Science 2.0â€™ or â€˜Open Scienceâ€™, and even that the long-promised â€˜knowledge economyâ€™ is actually coalescing, based on new products and services derived from the wealth of (open) data being created and made available.

â€˜The initial investment is scientific, but the ultimate return is economic and socialâ€™

John Wood, currently the Co-Chair of the global Research Data Alliance (RDA) as well as Chair of RDA-Europe, set out the case in his introduction to the Data Harvest report, and from the podium at the RECODE conference, that the new European commissioners and parliamentarians must first of all, not get in the way, and second, almost literally â€˜plan the harvestâ€™ for the economic benefits that the significant public investments in data, research and technical infrastructure are bringing.

The reportâ€™s irrepressible argument goes, â€œJust as the World Wide Web, with all its associated technologies and communications standards, evolved from a scientific network to an economic powerhouse, so we believe the storing, sharing and re-use of scientific data on a massive scale will stimulate great new sources of wealth.â€� The analogy is certainly helped by the fact that the WWW was invented at a research institute (CERN), by a researcher, for researchers. The web â€“ connecting 2 billion people, according to a McKinsey 2011 report, contributed more to GDP globally than energy or agriculture. The report doesnâ€™t shy away from reminding us and the politicians it targets, that it is the USA rather than Europe that has grabbed the lionâ€™s share of economic benefitâ€“ via Internet giants Google, Amazon, eBay, etc. – from the invention of the Web and that we would be foolish to let this happen again.

This may be a ruse to convince politicians to continue to pour investment into research and data infrastructure, but if so it is a compelling one. Still, the purpose of the RDA, with its 3,000 members from 96 countries is to further global scientific data sharing, not economies. The report documents what it considers to be a step-change in the nature of scientific endeavour, in discipline after discipline. The report – which is the successor to the 2010 report also chaired by Wood, “Riding the Wave: How Europe can gain from the rising tide of scientific data,” celebrates rather than fears the well-documented data deluge, stating,

â€œBut when data volumes rise so high, something strange and marvellous happens: the nature of science changes.â€�

The report gives examples of successful European collaborative data projects, mainly but not exclusively in the sciences, such as the following:

Lifewatch – monitors Europeâ€™s wetlands, providing a single point to collect information on migratory birds. Datasets created help to assess the impact of climate change and agricultural practices on biodiversity
Pharmacog – partnership of academic institutions and pharmaceutical companies to find promising compounds for Alzheimerâ€™s research to avoid expensive late-stage failures of drugs in development.
Human Brain Project – multidisciplinary initiative to collect and store data in a standardised and systematic way to facilitate modelling.
Clarin – integrating archival information from across Europe to make it discoverable and usable through a single portal regardless of language.

The benefits of open data, the report claims, extends to three main groups:

to citizens, who will benefit indirectly from new products and services and also be empowered to participate in civic society and scientific endeavour (e.g. citizen science);
to entrepeneurs, who can innovate based on new information that no one organisation has the money or expertise to exploit alone;
to researchers, for whom the free exchange of data will open up new research and career opportunities, allow crossing of boundaries of disciplines, institutions, countries, and languages, and whose status in society will be enhanced.

â€˜Open by Defaultâ€™

If the data harvest report lays out the argument for funding open data and open science, the RECODE policy recommendations focus on what the stakeholders can do to make it a reality. The project is fundamentally a research project which has been producing outputs such as disciplinary case studies in physics, health, bioengineering, environment and archaeology. The researchers have examined what they consider to be four grand challenges for data sharing.

Stakeholder values and ecosystems: the road towards open access is not perceived in the same way by those funding, creating, disseminating, curating and using data.
Legal and ethical concerns: unintended secondary uses, misappropriation and commercialization of research data, unequal distribution of scientific results and impacts on academic freedom.
Infrastructure and technology challenges: heterogeneity and interoperability; accessibility and discoverability; preservation and curation; quality and assessibility; security.
Institutional challenges: financial support, evaluating and maintaining the quality, value and trustworthiness of research data, training and awareness-raising on opportunities and limitations of open data.

RECODE gives overarching recommendations as well as stake-holder specific ones, a â€˜practical guide for developing policiesâ€™ with checklist for the four major stakeholder groups: funders, data managers, research institutions and publishers.

â€˜Open Changes Everythingâ€™

The Scottish government event was a pre-release of theÂ open data strategy, which is awaiting final ministerial approval, though in its final draft, following public consultation. The speakers made it clear that Scotland wants to be a leader in this area and drive culture change to achieve it. The policy is driven in part by the G8 countries’ â€œOpen Data Charterâ€� to act by the end of 2015 on a set of five basic principles â€“ for instance, that public data should be open to all â€œby defaultâ€� rather than only in special cases, and supported by UK initiatives such as the government-funded Open Data Institute and the grassroots Open Knowledge Foundation.

Improved governance (or public services) and â€˜unleashingâ€™ innovation in the economy are the two main themes of both the G8 charter and the Scotland strategy. The fact was not lost on the bureaucrats devising the strategy that public sector organisations have as much to gain as the public and businesses from better availability of government data.

The thorny issue of personal data is not overlooked in the strategy, and a number of important strides have been taken in Scotland by government and (University of Edinburgh) academics recently on both understanding the publicâ€™s attitudes, and devising governance strategies for important uses of personal data such as linking patient records with other government records for research.

According to Jane Morgan from the Digital Public Services Division of the Scottish Government, the goal is for citizens to feel ownership of their own data, while opening up â€œtrustworthy uses of data for public benefit.â€�

Tabitha Stringer, whose title might be properly translated as â€˜policy wonkâ€™ for open data, reiterated the three main reason for the government to embrace open data:

Transparency, accountability, supporting civic engagement
Designing and delivering public services (and increasingly digital services)
Basis for nnovation, supporting the economy via growth of products & services

â€˜Digital firstâ€™

The remainder of the day focused on the new EU Public Service Information directive and how it is being â€˜transposedâ€™ into UK legislation to be completed this year. In short, the Freedom of Information and other legislation is being built upon to require not just publication schemes but also asset lists with particular titles by government agencies. The effect of which, and the reason for the awareness raising workshop is that every government agency is to become a data publisher, and must learn how to manage their data not just for their own use but for public â€˜re-usersâ€™. Also, for the first time academic libraries and other â€˜cultural organisationsâ€™ are to be included in the rules, where there is a â€˜public taskâ€™ in their mission.

â€˜Digital firstâ€™ refers to the charging rules in which only marginal costs (not full recovery) may be passed on, and where information is digital the marginal cost is expected to be zero, so that the vast majority of data will be made freely available.

Robin Rice
EDINA and Data Library

Edinburgh supports trial data publication

Posted on August 18, 2014 by Robin Rice

I recently read in a Sense about Science tweet that a lone student asked the Principal of the University of Edinburgh if it would join the AllTrials Campaign and it became the first Scottish University to do so. Here’s his story – [Editor]

As a Nurse I frequently talk with patients about the treatments and medications they receive. These are often difficult conversations as it relies on the clinician having a library of background knowledge coupled with the most up-to-date data. Despite the wealth of knowledge that exists within the medical community there is an increasing body of research that highlights the large amount of clinical trial data that has gone unshared for many decades. This is the origin of the AllTrials campaign.

The best estimate is that around half of clinical trials have never been published. Recognising the need for change a group of academics founded AllTrials. AllTrials is an initiative headed by leading academic bodies such as the British Medical Journal, the Oxford Centre for Evidence-based Medicine and the Cochrane Collaboration. AllTrials calls for all past and present clinical trials to be registered and their full methods and summary results reported.

As a Nursing graduate of The University of Edinburgh and a current masters student within the Nursing department I felt I should engage with my University about the issue of clinical trial data sharing and about the AllTrials campaign. I wrote to Professor Sir Timothy Oâ€™Shea, the Principal of the University, who gave his support for AllTrials. As of July 2014 The University of Edinburgh became the first Scottish University to register its support for AllTrials. This move is inherently positive for Edinburgh, both as a global leader in health care and as an institution with a longstanding Edinburgh Data Library history. The campaign has had nearly 80,000 people sign its petition as well as just under 500 organisations register support.

This is an issue important issue for all of us. Show your support by signing the AllTrials petition.

Adam Lloyd
Nurse
Masters of Nursing in Clinical Research student
The University of Edinburgh

The views expressed are my own and do not reflect the views of The University of Edinburgh, the AllTrials campaign or any of its affiliates.

Data journals â€“ an open data story

Posted on June 6, 2014 by Pauline Ward

Here at the Data Library we have been thinking about how we can encourage our researchers who deposit their research data in DataShare to also submit these for peer review.

Why? We hope the impact of the research can be enhanced with the recognised added-value of peer review. Regardless whether there is a full-blown article to accompany the data.

We therefore decided recently to provide our depositors with a list of websites or organisations where they could do this.

I pulled a table together, from colleaguesâ€™ suggestions, from the PREPARDE project and the latest RDM textbook. And, very much in the Open Data spirit, I then threw the question open on Twitter:

â€œ[..]does anyone have an up-to-date list of journals providin peer review of datasets (without articles), other than PREPARDE? #opendataâ€�

â€¦and published the draft list for others to check or make comments on. This turned out to be a good move. The response from the Research Data Management community on Twitter was very heartening, and colleagues from across the globe provided some excellent enhancements for the list.

That process has given us confidence to remove the word â€˜Draftâ€™ from the title â€“ the list, this crowd-sourced resource, will need to be updated from time-to-time, but we are confident that weâ€™ve achieved reasonable coverage of the things we were looking for.

Another result of this search was the realisation that what we had gathered was in fact quite clearly a list of Data Journals. My colleague Robin Rice has now added a definition of that term to the list, and we will be providing all our depositors with a link to it:

https://www.wiki.ed.ac.uk/display/datashare/Sources+of+dataset+peer+review