The Power of Structured Data: Customizing Grid Views in LabKey Biologics

When analyzing biologics assay data, scientists often need to look beyond the results at related data to answer their research questions. Comparing lineage characteristics like which expression system was used to generate an experiment sample or details about the sample itself, such as the buffer used in it, can uncover crucial data patterns and insights.

This type of data exploration requires data to be captured in a structured manner and integrated into a central system where it can be easily accessed, queried, and analyzed.

Structuring Data for Maximum Value

LabKey Biologics provides tools to ensure that data is correctly structured and consistently stored. For each data type within the LabKey Biologics application, users are able to configure a specific structure, indicating the names of fields as well as their type. Because this data is consistently structured, a user can pull together relevant data from different sources for an integrated view of their data during analysis.

Integrating structured assay and sample data in LabKey Biologics using Sample IDs and look-ups.

For example, when looking at the results for a specific assay type, such as optical density, a user can add details about the samples themselves to the assay results data grid. This might include the buffer used, the expression system used to create it, or the name of the antibody (or other molecule) that was being produced.

Customized Data Views for Quicker Access

Users can customize the default assay data grid view to include these additional look-up columns. Customized default views provide quick access to all the data relevant to the user’s research, instead of having to join data each time they view the dataset. Users can also sort, filter, and search the data in these additional columns the same way they can with native assay data fields.[vc_cta h2=””]To see this functionality in action, request a demo! To learn more about LabKey Biologics check out our documentation and resources on the LabKey Support Portal. [/vc_cta]

3 Key Reasons Data Accessibility is Essential in Research

Modern research technologies have greatly increased the amount of scientific data being generated, but making full use of that data is still a major challenge. Data accessibility is a consideration at all stages of the research process; for bench scientists making data accessible to informaticians, for teams sharing data cross-departmentally, and for researchers making data accessible to the public.

The accessibility of data is essential for a number of reasons:

Data accessibility reduces duplication of experiments1. Minimizing Data Redundancy of Research Efforts

Research redundancy is a major problem within research organizations and across the research community. By making data accessible to their desired audience, researchers can reduce the number of redundant experiments conducted and instead iterate upon existing research to accelerate discovery.

Draw reliable conclusions from your experiment data2. Drawing More Reliable Conclusions from More Data

Broader data accessibility allows research teams to pool data and conduct analysis with greater confidence in their results. The more data a researcher has access to, the more statistical power they have to validate research conclusions and preempt questions of data quality.

Accessible data inspires novel approaches to answering scientific research questions.3. Inspiring Novel Questions from Different Approaches

New research questions are inspired by different research approaches and through the study of new methodology. Attacking scientific investigations from varying perspectives also helps reduce bias in analytics, experimental design, and conclusion drawing.

Expanding Data Accessibility with LabKey Server

LabKey Server not only helps teams collect and curate their data, but also helps make it accessible to collaborators and downstream researchers.

Web-Based Access

LabKey Server allows researchers to make their data accessible to a broad or narrow audience through a web-based portal. Web-based access makes it easy to share data as desired and allows interested collaborators to evaluate and alternatively analyze “self-serve” research data. This method of data sharing is both more secure than email (see fine grained permissions below) and much lower overhead than a standard institutional database as users can query, view, and export data without having to interface with a data scientist.

Fine-Grained Permissions

LabKey’s fine-grained permissions model makes secure, selective sharing of to data simple and reliable. With LabKey Server, teams can easily control who sees their data, restricting access to selected individuals or pre-defined groups, or making data accessible the general public. Researchers also have fine-grained control over what datasets are shared: either a single table of data or an entire research project.

Powerful Metadata

LabKey Server captures detailed metadata to help increase discoverability of research data and provide crucial context for other researchers who hope to explore, reanalyze, and/or expand upon it. Research teams can customize metadata captured for each of their data types and add organization specific metadata to support internal needs.

Interested in learning more about how LabKey Server can enhance the accessibility of your research data? Contact the LabKey team for more information or request a demo!

*To learn how configure accessibility features of LabKey Server, read documentation >

Two Key Things Your Spreadsheet-Based Research Data Management Strategy is Lacking

High-throughput analysis techniques are incredibly powerful and provide teams with more data than ever. While that depth of data often holds the key to scientific insights, organizing such large quantities of data in a consistent and discoverable way has become a major challenge for research teams.

Many teams rely on spreadsheet-based systems to organize and manage their data. This approach becomes less-effective as research scales because spreadsheet-based strategies lack two essential characteristics:

1. Consistency

Spreadsheet data management lacks consistencyManual file management relies on the individual contributor’s abilities to consistently create, name, and store data files. This opens the door to a wide range of human errors that will ultimately impact the discoverability and reliability of your data. Common consistency errors that result from manual data management include:

  • Poorly named files
  • Inconsistent locations
  • Duplicate files

2. Discoverability

Spreadsheet data management lacks discoverabilityCollecting data is a giant hurdle in research, but in reality, it is just the first of many. Researchers need to be able to locate datasets of interest in order to conduct analysis. In a file based environment, discoverability of files is dependent on the consistency with which they are maintained. Were they saved in the correct location? Have they been named according to an agreed upon convention? Is there a clear authoritative file or are there duplicates?

A hitch in any one of these areas can severely hinder the discoverability of your data and make it significantly more difficult to:

  • Track what research data has already been collected
  • Find the data you are looking for when it comes time to analyze

Biology-Aware Data Management with LabKey Server

Scientific data management systems like LabKey Server, help increase the consistency and discoverability of your research data. LabKey Server increases the consistency of data management by providing structured data grids for storing various type of research data. Each data grid type also captures relevant metadata, specific to that data type, in order to help make data more discoverable.

Research-Centric Data Structures

Unlike spreadsheets that treat all types of data the same, LabKey Server provides four primary data structures with unique features to better support common types of research data.*

LabKey Assays – Assay data grids capture data generated from individual experiment runs. Assay data is automatically structured in a batch-run-results hierarchy when data files are added. LabKey Server supports data a variety of common assay designs out of the box, but teams can also design their own assay data structure using LabKey’s General Purpose Assay Design.

LabKey Datasets – Datasets track patient/subject measurements over time. LabKey datasets are automatically aligned and joined together, making it easy to query the integrated data and to create visualizations from multiple datasets.

LabKey Specimens – Specimen repositories track the status of each specimen and vial in your inventory. Built-in reports provide a birds-eye view of specimen information, and advanced search capabilities allow for easy location of specimens.

LabKey Lists – LabKey lists provide general purpose, online, interactive grids for any tabular data. Data stored as a LabKey list can be sorted, filtered, and visualized using built-in tools.

Storing data in a consistent, structured manner is the key for teams that hope to achieve maximum efficiency in operations and maximum value from their data. Not only is it much simpler to find data when it is stored in an expected location, but the centralization and integration makes it possible to query data to more quickly locate information of interest.

Interested in learning more about how LabKey Server can increase consistency and discoverability of your research? Contact the LabKey team for more information or request a demo!

*To learn how to add data to LabKey Server, read documentation >

Genomics England and LabKey: Creating and securing “a dialogue between the clinical context and researchers.”

Genomics England 100,000 Genomes ProjectIn late 2015, Genomics England began working with LabKey to develop a LabKey Server-based data management and exploration portal that would facilitate the knowledge sharing dialogue between clinicians and researchers as part of the UK’s 100,000 Genomes Project.

The 100,000 Genomes Project, as characterized by Genomics England’s Chief Technology Officer, Jim Davies, is intended to promote “a dialogue between the clinical context and researchers.” This project, the largest national sequencing project of its kind in the world, will give both clinicians and researchers access to an unprecedented depth of data and information through the sequencing of 100,000 genomes from approximately 70,000 people including patients with a rare disease and their families, as well as patients with cancer. The mission of this dialogue is to ultimately bring benefit to patients and to enable new scientific discovery and medical insights in an ethical and transparent manner that will promote the development of a UK genomics industry.

The first phase of this collaboration centered around providing clinicians and researchers access to centralized phenotypic and sample information gathered from sites across the UK while ensuring security and privacy of patient information. The LabKey team focused development efforts during phase 1 on the aggregation, review, and integration of phenotype and genotype information from cancer and rare disease patient data.

This phase has provided:

  • Secure, extensible and reliable pipelines for data collection from 13 NHS Genomic Medicine Centres leading participant recruitment and partnering hospitals & clinics
  • Medical review workflow to assess data gathered of participants and families
  • Web portals for secure access of data
  • LabKey Server’s built-in reporting, querying, and visualization tools

LabKey and Genomics England are bringing the value of dialogue to the development process. LabKey is excited to continue its tradition of collaborating closely with its partners. With Genomics England, LabKey has developed a deep shared understanding of goals leading to a phased development roadmap of the LabKey Server platform. Close communication and collaboration will enable LabKey to flexibly accommodate new requirements and priority shifts to ensure a high-quality final product.

LabKey looks forward to the next phase of development of a “research” LabKey Server platform that will securely store and enable access to de-identified information that can be used by clinicians and researchers for advanced analysis. LabKey is proud to support Genomics England’s mission to promote a productive “dialogue” that will improve patient outcomes and scientific progress.

To learn more about the 100,000 Genomes Project, visit: https://www.genomicsengland.co.uk/the-100000-genomes-project/

Allen Institute for Cell Science Uses LabKey to Simplify Workflows and Accelerate Target Identification

LabKey partners at the Allen Institute for Cell Science are doing things a little differently. Launched in 2014 with a contribution from founder and philanthropist Paul G. Allen, the Institute was formed to integrate diverse technologies and approaches to study the cell as an integrated system. They are asking fundamental questions about cellular behavior to better understand healthy and pathological cells. Data and tools developed by the Institute are made publicly available to researchers around the world.

The Allen Institute for Cell Science team uses genome editing to add fluorescent markers to proteins in key cellular machinery and uses light microscopy to study the organization of that machinery and how it changes. During the generation of genome edited cell lines, they conduct quality control steps, including genotyping, stem cell marker analysis, karyotype testing, deep sequencing, and image-based assays. They also generate clonal lines for each gene and ensure that the cell line is useful for long term experiments. This work requires tracking cells, cell lines, genes, and all other quality control components through approximately 40 workflow steps.

The Allen Institute for Cell Science uses LabKey Server to capture metadata and assay data on each gene, cell, cell line, clone, and vector at each stage in the workflow: information that was previously housed in spreadsheets and raw data files. The Institute relies on the integrated views and query-ability of their LabKey managed data to efficiently identify target cells/cell lines to explore. Their use of LabKey Server also allows them to track the status of each entity as it makes its way through the processing pipeline, as well as monitor their complete workflow.

Allen Institute for Cell Science + LabKey Server Workflow

By centralizing their data, the Allen Institute for Cell Science will be able to accelerate their analysis and use insights gathered about their workflow to optimize their operations.

[vc_cta h2=”” shape=”square” style=”custom” custom_background=”#ededed”]Interested in using LabKey Server to optimize your workflow and accelerate data analysis?  Contact the LabKey team for more information![/vc_cta]

Next Generation Clinical Trial Transparency: Providing Management and Analysis of RNA Sequencing Data

LabKey has begun work with long-time partner the Immune Tolerance Network (ITN) to further clinical trial transparency by extending the LabKey Server platform to support management of RNA sequencing (RNA-Seq) and Next Gen Sequencing data. The ITN will be providing access to sequencing data, connecting raw data to downstream sequence analysis and visualizations that incorporate clinical data and other assay data through ITN’s research web portal, ITN TrialShare.

ITN TrialShare is an application built on the LabKey Server platform that supports both operational data management and post-publication sharing of data from clinical trials. ITN TrialShare was the first application of its kind to allow direct linking from major clinical trial publications to de-identified participant-level data and analyses.

As part of this new project, the LabKey team will provide custom development and configuration services to extend ITN TrialShare to support integration of and enhanced access to sequencing data. Improvements being explored as part of this project include:

  • Support for big data downloads such as FAST-Q and BAM raw data files
  • Integration with sequencing data transfer tools such as Globus Genomics
  • The addition of assay data filtering for RNA-Seq data, T-cell Repertoire, and WGS data
  • Support for Shiny in LabKey’s reporting framework, allowing real-time visualization of big data

[vc_cta h2=”” shape=”square” style=”custom” custom_background=”#ededed”]Interested in using LabKey Server to manage your research and help promote data transparency at your organization?  Contact the LabKey team for more information![/vc_cta]

About the Immune Tolerance Network

The Immune Tolerance Network (ITN) is a collaborative network for clinical research focused on the development of therapeutic approaches for asthma and allergy, autoimmune diseases, type 1 diabetes and solid organ transplantation that lead to immune tolerance. The ITN encompasses over 250 clinical sites and investigators and 14 core labs supporting 95 clinical trials (23 Allergy trials; 43 Autoimmunity trials; 29 Transplant trials). Read about LabKey and ITN’s partnership to develop ITN TrialShare >

LabKey Kicks Off Abstraction and NLP Pipeline Project for NCI SEER

In support of the National Cancer Institute (NCI) and the Department of Energy (DoE) initiative to use large-scale computing to influence cancer science, NCI’s Surveillance, Epidemiology, and End Results (SEER) Program has partnered with LabKey to develop an abstraction workflow and Natural Language Processing (NLP) pipeline that will automate the annotation and review of free-text pathology reports.

The NCI SEER Program works to provide information on cancer statistics in an effort to reduce the burden of cancer among the U.S. Population. SEER currently collects and publishes cancer incidence and survival data from population-based cancer registries covering approximately 30 percent of the U.S. population. The registries receive at least one unstructured pathology report on the more than 450,000 cases reported annually that are used in conjunction with other sources to abstract relevant information on the cases.

The initial version of the application will allow SEER to identify and select pathology reports of interest using a Linguamatics-based text-mining tool and make them available in a LabKey Server portal for manual annotation and abstraction of key elements. An annotation and task management pipeline will manage the stages of the abstraction, annotation, and review process, ensuring consistency by standardizing tasks and automating the workflow. To ensure the security of data being processed and generated, the application will utilize LabKey Server’s role-based security model and facilities that ensure regulatory compliance.

Future phases of development will introduce the use of Natural Language Processing engines to further accelerate the annotation process. Manually abstracted data from the initial phase of the project will be used by DoE laboratories to help develop and train NLP algorithms that will then be used to automate the abstraction of large numbers of pathology reports.


LabKey Server for Your Research

Interested in exploring LabKey Server as a bioinformatics solution for your organization? Contact us and the LabKey team will work with you to understand your needs and help determine if LabKey is the right fit for your research.

[wpi_designer_button text=’Contact Us’ link=’/about/contact-us/’ style_id=’75523′ target=’self’]

O’Connor Lab Applauded for Real-Time Data Sharing

Our collaborators at the O’Connor Lab (University of Wisconsin-Madison) are making headlines for releasing real-time data via LabKey Server to help accelerate Zika virus research. Recently featured in the Nature article “Zika researchers release real-time data on viral infection study in monkeys,” Dave O’Connor and his team are being applauded for making their research available so quickly.

“O’Connor’s team is to be lauded for their efforts to make their Zika virus data publicly available as soon as possible,” says Nathan Yozwiak, a senior scientist in Pardis Sabeti’s laboratory [computational geneticist at the Broad Institute and Harvard University in Cambridge] “Distributing up-to-date information — in this case, animal model data — as widely and openly as possible is critical during emergencies such as Zika, where relatively little is known about its pathogenesis, yet public concerns and attention are so high.”

The O’Connor lab uses LabKey Server to manage their extensive list of experiments, Illumina sequencing data, purchases, oligonucleotides, freezer samples, and other lab inventory, as well as to provide basic electronic lab notebook (ELN) functionality. To make their LabKey Server data public, the O’Connor lab simply had to update the study permissions.

“It was easy for the ZEST members to make their online lab notebook open to all, O’Connor says. The team uses the biomedical-research collaboration system LabKey Server, as does the Wisconsin National Primate Research Center in Madison, which is where many of the ZEST collaborators work and which (along with the US National Institutes of Health) is supporting the research. Researchers created a study to store and update their data, and simply had to switch permissions to allow anyone to view it. Meanwhile, regulatory agencies at the University of Wisconsin–Madison understood that the work was time-sensitive and expedited approvals for animal care and biosafety (without reducing scrutiny, O’Connor adds).”

[wpi_designer_button text=’Read the Full Article on Nature.com’ link=’http://www.nature.com/news/zika-researchers-release-real-time-data-on-viral-infection-study-in-monkeys-1.19438′ style_id=’75523′ target=’self’]