New R Tools Simplify FAIR Data Sharing for WASH Researchers
A new suite of R tools aims to make research data sharing more accessible for scientists working in water, sanitation, and hygiene (WASH) fields. The packages, including 'washr' and 'fairenough', streamline the process of creating and publishing FAIR-compliant data packages, lowering technical barriers and encouraging open science practices.
Researchers in the water, sanitation, and hygiene (WASH) sector often generate valuable datasets that remain underutilized, partly due to limited adoption of FAIR (Findable, Accessible, Interoperable, Reusable) data practices. This challenge is pronounced among teams working in resource-limited environments, where technical expertise and tools for data sharing may be lacking.
To address these issues, the Open Science initiative 'openwashdata', operating through the GHE network, conducted surveys among WASH researchers. The findings confirmed that many researchers still rely on data storage methods that hinder data portability and interoperability, and that programming experience—particularly with the R language—varies widely across the community. These barriers restrict the sharing and reuse of valuable research data.
In response, the team developed two open-source R packages: 'washr' and its successor, 'fairenough'. The 'washr' package streamlines the creation of publication-ready data packages by leveraging R's development tools. However, recognizing ongoing obstacles for non-expert users, the newer 'fairenough' package further reduces technical complexity. It introduces a one-command pipeline, guiding users through every step of data package creation, website generation, and publication with minimal manual input.
A standout feature of 'fairenough' is its integration of large language models (LLMs) via the 'ellmer' library to automatically generate data dictionaries, further reducing the need for specialized knowledge. The package also automates metadata creation, ensures proper documentation, implements version control through Git and GitHub, and enables Digital Object Identifier (DOI) assignment using Zenodo, supporting each aspect of the FAIR data principles.
To help researchers adopt these tools, the team produced a comprehensive publishing guide using R and Quarto. This resource offers step-by-step instructions for creating reproducible data packages and websites, making datasets available for download in widely used formats such as CSV and XLSX.
The 'fairenough' package debuted to positive feedback during a lightning talk at the LatinR 2025 conference, where its ability to streamline data publication processes for researchers—particularly those from Spanish and Portuguese-speaking communities—was well received. The developers see improving knowledge and adoption of open science practices as key to maximising the global impact and visibility of WASH research.
For researchers and practitioners looking to make their data more open and reusable, resources and guides for getting started with 'fairenough' are available online.
Reference: r-bloggers.com
Related Posts
E.ON Modernises Energy Grid with SAP S/4HANA and AI
E.ON is leveraging SAP S/4HANA to standardise grid data, streamline infrastructure, and enable AI-powered applications such as predictive maintenance and customer automation. The company is focusing on internal technical capabilities, cybersecurity, and embedding digital tools directly into core operations to support reliability and growth in the energy sector.
Walmart Limits Employee AI Use to Manage Rising Costs
Walmart has imposed limits on employee use of its internal AI assistant, Code Puppy, in response to unexpectedly high costs associated with large language model (LLM) usage. The move highlights broader challenges faced by large enterprises as AI billing models shift from flat-rate subscriptions to usage-based pricing.
NHL Modernises Media Operations with VAST Data Platform
The National Hockey League (NHL) has revamped its media storage and distribution systems through a multi-year partnership with VAST Data. The initiative replaces legacy archives and in-arena storage, enabling faster, more efficient media workflows and paving the way for advanced analytics. This modernisation is expected to enhance fan experiences and streamline media operations across the league.