Data assumptions for the eLignin Microbial database
The data in eLignin broadly falls into two categories:
1. Reported empirical data from scientific publication
2. Inferred but unverified data predicted from the empirical data
1. Reported empirical data from scientific publication
The empirical section of the pages in eLignin are strictly based on empirical evidence reported in the scientific literature. Even if there is obvious inference — e.g. a strain has been shown to contain a pathway that implies metabolism of certain compounds — such connections will not be reported unless there is evidence for that particular experiment in the literature. Inference is handled separately in differently marked sections on the pages in the database (described in Section 2 below).
Empirical data sections in eLignin are marked with: Empirical . This badge indicates that the data in that section is directly supported by primary literature and not inferred.
Utilizable substrates are based on physiological experiment that have investigated if a microbial strain can consume a given aromatic compound during cultivation in a suitable growth medium. It does not mean that the aromatic compound can sustain growth as the sole carbon source.
Pathways are typically reported based on genomic analysis and possibly also by in vitro enzyme assays. For the empirical data, eLignin will not assume any link between an organism's utilizable substrates, pathways and enzymes unless there is empirical evidence for these. It will try to predict it in the inference section of the pages.
Based on the above constraints, the following relationships gaps are possible in the empirical data in eLignin:
- Organisms with substrates but no pathway
- Organisms with pathways but no enzyme or substrates
- Substrates not in any reaction (not all substrates have known degradation pathways, especially for lignin macropolymers )
- Reactions with no enzyme link (spontaneus reactions)
2. Inferred but unverified data predicted from the empirical data
eLignin uses the empirical data as a starting point to infer additional relationships that have not been experimentally verified. These predictions are clearly labelled throughout the database and should be treated with caution. All inferences follow a defined logical chain described below.
Predicted data sections in eLignin are marked with: Predicted . This badge indicates that the data in that section is computationally inferred from relationships in the database and not experimentally verified.
2.1 Predicted pathway membership for organisms
Inference chain: Utilizable substrate → pathway
If an organism is recorded as being able to utilise a compound that enters a metabolic pathway as a reactant, the organism is predicted to carry that pathway. This prediction is only made if the organism is not already confirmed for that pathway by empirical evidence.
Shown on: organism detail pages, pathway detail pages, substrate detail pages.
2.2 Predicted enzyme presence for organisms
Inference chain: Confirmed pathway → enzyme known to be involved in that pathway
If an organism is empirically confirmed to carry a pathway, and that pathway is known to involve a specific enzyme (in other organisms), the enzyme is predicted to be present in the organism. This prediction is only made if the organism is not already confirmed to possess the enzyme. The pathway — and the reference confirming the organism carries that pathway — are shown as the basis for each prediction.
This directly addresses the known data gap of "organisms with pathways but no enzyme" in the empirical records.
Shown on: organism detail pages, enzyme detail pages, pathway detail pages.
General caveats
- All predicted relationships are computationally inferred and have not been independently verified in the laboratory.
- The absence of a predicted relationship does not mean the relationship does not exist — it only means the inference cannot be made from the current data.
- Predictions are updated automatically as new empirical data are added to the database.