Showing posts with label science. Show all posts
Emerging Concepts about the Role of Protein Motion in Enzyme Catalysis
Corresponding authors:
Sharon Hammes-Schiffer University of Illinois, Urbana−Champaign
Judith Klinman, University of California, Berkeley
Acc. Chem. Res., 2015, 48 (4), pp 899–899
DOI: 10.1021/acs.accounts.5b00113
More about authors:
The Hammers-Schiffer Research Group: http://hammes-schiffer-group.org/
The Klinman Group: http://www.cchem.berkeley.edu/jukgrp/klinman_group/Home.html
The Special Issue of Accounts of Chemical Research on Protein Motion in Catalysis addresses one of the most active and compelling areas of investigation regarding protein function: the relationship of motions within a protein to catalytic rate enhancement, allosteric control, and thermal adaptation. There is growing acceptance that our ability to understand and predict protein function must go beyond the static views obtained from traditional X-ray crystallography and incorporate both local and global protein conformational landscapes that can involve large segments of a protein and take place on time scales from femtosecond to millisecond. The fields of computation and experimentation, which at times have appeared in conflict, are increasingly providing synergistic pictures of protein behavior.
The most direct properties that can be studied are for the ground state protein alone or in complex with substrate or modulator ligands. The emerging importance of Markov state models and the computation or measurement of side chain entropy is facilitating the evaluation of interconverting ground state structures and the relative importance of surface versus buried protein motions. Advances in NMR have provided the ability to detect low population protein substates, as do room temperature Ringer X-ray studies. Monitoring the progression of the enzyme–substrate complexes toward the activated complex introduces additional challenges for achieving sufficient spatial and temporal resolution of functionally linked protein motions. A wide range of studies has supported the concept of modulation of reaction barrier width and height, as well as active site electrostatics, through protein motions. What is particularly exciting is that this concept, which originally emerged from the properties of room temperature electron and most recently hydrogen tunneling, has become apparent in many different classes of enzyme reactions.
Advances in methodology are critical to the growth within this field of inquiry. Well-established computational approaches that include mixed quantum mechanical/molecular mechanical (QM/MM) approaches, empirical valence bond (EVB) potentials, molecular dynamics free energy simulations, and transition path sampling methods are being expanded to address challenges such as sampling for much longer (i.e., microsecond) time scales and incorporating a growing number of atoms in the region treated quantum mechanically. The experimental toolkit has also gained traction, with the availability of techniques that can be monitored and evaluated as a function of perturbations that alter catalytic efficiency. In the picosecond to nanosecond time regime, changes to vibrational frequencies can be observed through the selective modification of amino acid side chains or the use of bound substrates with vibrational frequencies that lie outside of the protein envelope. Time-resolved fluorescence measurements continue to provide important insight into local motions that control the lifetime and emission wavelengths for appropriately placed chromophores. On the much longer time scale, seconds to hours, hydrogen–deuterium exchange is a well-validated tool to evaluate changes in local protein unfolding and its dependence on perturbations to the protein or environment. One critically needed time scale lies within the microsecond to millisecond regime, which is the one most generally implicated for the real time interconversion among multiple protein substates. Future methodological developments within this time regime will be a boon to the field of protein dynamics.
Unraveling the mysteries of protein motions and conformational sampling is also relevant to protein design efforts. Attempts at rational protein design have often focused on the structural aspects of ligand binding and enzyme catalysis. The recent discoveries and insights summarized in this special issue suggest that protein motion and conformational sampling should also be taken into account in protein design strategies. Thus, these concepts could have broad implications for drug design and for the design of more effective catalysts for biomedical and technological purposes.
The electronic version of this article is the complete one and can be found online at: http://pubs.acs.org/doi/full/10.1021/acs.accounts.5b00113
Image source: allacronyms.com, ACS
Imaging technologies from bench to bedside
Corresponding author: Ravinder Reddy krr@mail.med.upenn.edu
Author Affiliations
Center for Magnetic Resonance and Optical Imaging, Perelman School of Medicine, Department of Radiology, University of Pennsylvania
Journal of Translational Medicine 2015, 13:97 doi:10.1186/s12967-015-0449-5
More about author
Editorial
The last few decades have seen tremendous advances in medicine that have enhanced understanding of pathophysiological processes at the cellular and molecular level, and led to the development of increasingly sophisticated diagnostic imaging technologies. Early detection of disease induced molecular and functional changes before induction of irreversible structural changes is key for optimal treatment efficacy. Non-invasive imaging modalities, such as positron emission tomography (PET) [1], single photon emission computed tomography (SPECT) [2], computed tomography (CT) [3], optical tomographic technologies [4], magnetic resonance imaging (MRI) [5], ultrasound (US) [6], and X-rays play a vital role in both the diagnosis and monitoring of disease in response to therapy. These techniques cover a broad range of spatio-temporal resolution and varying degrees of sensitivity and specificity to different molecular changes, and in many cases provide complementary information [7],[8]. Recently discovered molecular targets of various disease states, including oncology, neurodegenerative and neuropsychiatric, cardiovascular, and musculoskeletal pathologies, drive further developments in the imaging field to detect these new molecular markers. Ultimately, these technologies contribute to improved disease management and personalized patient care.
Standard-of-care medical imaging techniques such as X-rays, US, CT and MRI provide exquisite structural details of human anatomy. These methods are the first-line techniques in clinic for diagnosis and characterization of disease, based primarily on structure/morphology such as size, texture and tissue attenuation [8]. In addition to providing diagnostic information, the US modality has the additional benefit of use as a therapeutic tool [6],[7],[9].
Functional nuclear medicine techniques (PET and SPECT) provide a unique, non-invasive assessment of intracellular processes and enzyme trafficking, receptors and gene expression, and serve as the underpinnings of molecular medicine. These techniques provide non-invasive diagnostic information about biochemical and physiological process ranging from glucose metabolism to gene expression by evaluating the kinetics of short-lived radioisotope tracers. While many promising tracers have been synthesized that target a variety of metabolic pathways or specific markers 18F-fluorodeoxyglucose (FDG), a glucose analogue is the main radiotracer in clinical practice today. In addition, these functional nuclear medicine techniques are also being used in research and clinical settings to detect and evaluate Alzheimer’s disease, metabolic viability of cardiac tissues, in vivo gene expression, and in tracking of cancer metastasis to different organs [7],[8],[10]-[12].
MRI is one of the most powerful and versatile non-invasive techniques. The major advantage of MRI is that it provides high-resolution, three-dimensional images of tissue structure, as well as functional and metabolic information. Furthermore, MRI is performed in vivo without the use of any ionizing radiation, allowing for repeated study. Several advanced MRI methods have been introduced to monitor the structural [13], functional [14],[15] as well as biochemical changes in various diseases. Magnetic resonance spectroscopy (MRS), which provides the information about the biochemical signatures, is an additional important clinical research tool to assess and characterize disease pathophysiology [16],[17].
Using MRS, enriched metabolites (e.g. 13C enriched) can be used to probe endogenous reaction kinetics. Latest advances in chemical exchange saturation transfer (CEST) MRI show promise in detecting several endogenous metabolites and proteins with substantially enhanced sensitivity (at least an order of magnitude) compared to conventional MRS [18]-[25]. Recent developments in hyperpolarized imaging based on dynamic nuclear polarization (DNP) of 13C enriched pyruvate are yielding highly promising preclinical results [26],[27] exploring in vivo reactions in oncology and other disease conditions. Some very preliminary results showing the promise of these methods in addressing clinical problems in patients have been demonstrated [28].
Optical imaging is another emerging imaging modality with high potential for improving diseases diagnosis and treatment, which can be readily set up at the patient’s bedside or in the operating room [4],[29]-[31]. Optical imaging uses non-ionizing radiation and offers potentially to image organs, tissues as well as smaller structures including cells and molecules using their unique photon absorption or scattering profiles. It also differentiates between native soft tissue and tissue labeled with endogenous or exogenous probes based on their wavelength dependent photon absorption or scattering pattern [32]-[35]. Despite limitations in their spatial resolution, optical imaging methods offer capabilities for studying functional and molecular events in different pathophysiological conditions. There are several techniques in optical imaging that are currently being used both in research and clinical setting for evaluating various diseases and therapeutic responses [30],[36]-[38]. Potentially, optical imaging can also be combined with other imaging modalities to improve the patient’s clinical management.
PET, SPECT and near-infrared reflectance fluorescence optical imaging techniques have relatively high sensitivity and can detect compounds with concentrations in micro- to pico-molar range [7]. Despite the high sensitivity these methods are beset by a relatively low spatial resolution (5 to 10 mm in clinical setting). Also, in many cases the emitting ligands may lose the specificity. One issue with nuclear medicine techniques is the use of nuclear radiation, which precludes their repeat use in short time spans. On the other hand, MRI provides high spatial resolution (in hundreds of micrometers range), but is relatively insensitive, in comparison to nuclear medicine techniques mentioned above; it requires concentrations of metabolites to be detected to be in the millimolar range and few endogenous molecules or metabolites can be imaged [39].
A milestone in the field of diagnostic imaging is the emergence of integrated structural and functional modalities such as combined PET-CT and PET-MRI [40],[41]. These integrated modalities provide concurrent structural, molecular and functional information, improve the multimodal imaging correlations and ease the patient burden for multiple imaging sessions.
Combining these advanced imaging techniques will result in improved precision of the data that are intrinsically more sensitive to the underlying pathophysiology than the morphological features available in routine structural imaging. Over the years, all these powerful imaging techniques have been improving the way the diseases are diagnosed, therapeutic responses are monitored and dramatically enhancing the practice of medicine making it more prognostic, preventative and personalized. Despite these advances, many technological innovations in these imaging modalities are still in research setting. Transferring these technologies into clinical setting requires an intense collaborative effort between researchers in imaging physics, instrumentation, image processing, biologists, chemists, and regulatory bodies as well as clinicians from all branches of medicine. This journal section facilitates communication of advances in translating the imaging modalities from mere research tools to clinical setting. We welcome research articles from all the stakeholders in this field.
The electronic version of this article is the complete one and can be found online at: http://www.translational-medicine.com/content/13/1/97
Image source: National Institute of Biomedical Imaging and Bioengineering
Still 0 and 1? Check this out: Memory leads the way to better computing
Correspondence: H.-S. Philip Wong & Sayeef Salahuddin
H.-S. Philip Wong - Department of Electrical Engineering and the Stanford SystemX Alliance, Stanford University, Stanford, California 94305, USA
Sayeef Salahuddin - Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, California 94720, USA
Nature Nanotechnology 10, 191–194 (2015) doi:10.1038/nnano.2015.29
More about Authors:
http://web.stanford.edu/~hspwong/
https://www.eecs.berkeley.edu/Faculty/Homepages/salahuddin.html
Introduction
Current memory devices store information in the charge state of a capacitor; the presence or absence of charges represents logic 1's or 0's. Several technologies are emerging to build memory devices in which other mechanisms are used for information storage. They may allow the monolithic integration of memories and computation units in three-dimensional chips for future computing systems1. Among those promising candidates are spin-transfer-torque magnetic random access memory (STT-MRAM) devices, which store information in the magnetization of a nanoscale magnet. Other candidates that are approaching commercialization include phase change memory (PCM), metal oxide resistive random access memory (RRAM) and conductive bridge random access memory (CBRAM).
Today's computing systems use a hierarchy of volatile and non-volatile data storage devices to achieve an optimal trade-off between cost and performance2. The portion of the memory that is the closest to the processor core is accessed frequently, and therefore it requires the fastest operation speed possible; it is also the most expensive memory because of the large chip area required. Other levels in the memory hierarchy are optimized for storage capacity and speed (Fig. 1). The main memory is often located in a separate chip because it is fabricated with a different technology from that of the microprocessor.
For over 30 years, static random access memory (SRAM)3 and dynamic random access memory (DRAM)3 have been the workhorses of this memory hierarchy4. Both SRAM and DRAM are volatile memories — that is, they lose the stored information once the power is cut off. For non-volatile data storage, magnetic hard disk drives (HDDs) have been in use for over five decades5, 6, 7. Since the advent of portable electronic devices such as music players and mobile phones, however, solid-state non-volatile memory known as Flash memory8 has been introduced into the information storage hierarchy between the DRAM and the HDD. Flash has become the dominant data storage device for mobile electronics; increasingly, even enterprise-scale computing systems and cloud data storage systems are using Flash to complement the storage capabilities of HDD.
For complete FREE article for:
- Electrical manipulation of magnetism in ferromagnets
- Electrical control of magnetization in multiferroics
image source: Controlling size and helicity of a spin-vortex skyrmion
Tips: For all Scientists: Starting up a career
Start up a career:
Tips extracted from the articles:
Lesson 1: Write a business plan.
Lesson 2: Get the right patent.
Lesson 3: Enter a contest.
Lesson 4: Funding comes in many forms.
Lesson 5: The science isn't everything
Lesson 6: Rent a bench.
Lesson 7: Assemble a good team.
Lesson 8: Biotech companies can be virtual.
Lesson 9: Entrepreneurship = experience
You will find the technology you are developing is only a tiny factor in your entrepreneurial journey.
image source: 10 Entrepreneur Lessons Not Taught in Classroom
Open questions: seeking a holistic approach for mitochondrial research
Correspondence: Heidi M McBride heidi.mcbride@mcgill.ca
Montreal Neurological Institute, McGill University, 3801 University Avenue, Rm 622C H3A 2B4, Montreal H3A 0G4, QC, Canada
BMC Biology 2015, 13:8 doi:10.1186/s12915-015-0120-x
[More about author]
Collaborate
Most of us are not geniuses, and cannot operate with an encyclopaedic knowledge of metabolism, calcium homeostasis, tissue physiology, bioenergetics and lipid chemistry. On the other hand, clinician scientists or physiologists who hope to incorporate mitochondria into their signalling paradigms feel overwhelmed with the complexity of the organelle, the experimental approaches, and sometimes, the dogma common to such established fields. The first step is an obvious one - forge meaningful collaborations that will truly push the field forward. I think we are finally past the point where non-mitochondrial scientists simply write us off, assuming that the mitochondria are a known entity, uninteresting, boring. Indeed the potential for fundamental new concepts in mitochondrial function has never been higher, and the disease relevance is clear. Mitochondrial pathways are practically untouched as a therapeutic target, for example.
For those of us working on the fundamental aspects of mitochondrial function, we must work harder to consider the physiology of real tissues. I’m not suggesting we abandon our fundamental projects, certainly not! Basic discoveries will remain the lifeblood of clinical development. But with collaborations we can extend our studies simultaneously and move into ‘real’ cells. Adapting these models will more rapidly push our discoveries up the ladder of biomedical translation. My own collaborations have provided me with confidence and helped me to understand complex physiologies that would otherwise have not crossed my radar screen. It sounds obvious, but funding agencies and promotion mechanisms do not always reward collaborations enough. Grants need a single principal applicant and team grants can be more political than functional. It is also clear that collaborations are more difficult than simply continuing along a successful, independent track. However, understanding the complexities of mitochondrial function and signalling will require open, and sometimes challenging, collaborations.
The electronic version of this article is the complete one and can be found online at: http://www.biomedcentral.com/1741-7007/13/8
There is an interesting video on Youtube about mitochondria!! - "Power Pack - The Mitochondria Rock Song"
image source: Blame it On Your Mother
Nanyang Technological University produced a 3D printed concept car
Of course this is not the first 3D printer car, but it is good to see a group of students from Nanyang Technological University to get it done.
They are NTU Venture 8 & NTU Venture 9, both with round 150 parts are 3D-pritned.
For the NTU Venture 8, honeycomb structure and a unique joint design was employed, to ensure the car doesn't come apart. Also, the car chassis is see-through when hit by the light at the right angle. Presumably, most of the light will be absorbed by the solar cells so that the inside of the car doesn't heat up overmuch during summer.
NTU Venture (NV) 9, a slick three-wheeled racer which can take sharp corners with little loss in speed due to its unique tilting ability inspired by motorcycle racing.
NTU unveils Singapore’s first 3D-printed concept car
3D-printed green vehicle to blaze a trail for future car technology
This first 3D printed car is like this:
image source: straitstimes
Valorisation of food waste to biofuel: current trends and technological challenges
Corresponding author: Carol SK Lin carollin@cityu.edu.hk
School of Energy and Environment, City University of Hong Kong, Tat Chee Avenue, Kowloon, Hong Kong
Sustainable Chemical Processes 2014, 2:22 doi:10.1186/s40508-014-0022-1
<More about author>
Introduction
Food waste is creating serious environmental and social problems in Hong Kong and across the world. According to a recent report released by the Hong Kong Environment Bureau, among the 9,000 tonnes of municipal solid waste that is thrown away everyday at landfills, 40% of which is composed of “putrescibles” [1]. These putrescibles are organic wastes that are known to create odour upon decomposition. Approximately, 90% of putrescibles are food wastes. Food waste can be raw, cooked, edible and inedible parts generated during production, storage distribution, and consumption of food stuffs. In 2011, Hong Kongers threw away approximately 3,600 tonnes of food waste everyday [1],[2]. Two third of the food waste was obtained from household; whereas, one third of food waste came from commercial and industrial sources. Hong Kong is not the only country generating large quantities of food waste. For instance, other developed cities like Taipei and Seoul are producing 182,000 tonnes/per year and 767,000 tonnes/per year of food wastes respectively [1],[2].
The food wastes produced in Hong Kong includes rotten fruits, vegetables, fish, poultry organs, fruits and vegetable peelings, meat, fish, shellfish shells, bones, food fats, sauces, condiments, soup pulp, Chinease medicinal pulp, egg shells, cheeses, ice cream, yogurts, tea leaves, teabags, coffee grounds, breads, cakes, biscuits, desserts, jam, different cereals, leftover of cooked food, BBQ raw or cooked leftovers, and pet food [1]. The Hong Kong Government plans to cut down the food waste that goes to landfill to approximately 40% by 2022. Landfills are the most common place for garbage deposition. Landfills spread offensive smell and are known to cause hazardous effects on people, animals, and the environment. Landfills are unsustainable as they produce methane which is a common green house gas. Furthermore, landfills also generate large amount of harmful leachate when rainwater falls on the garbage. This leachate can contaminate water and soil. Nevertheless, anaerobic digestion of food wastes that occurs naturally in the absence of oxygen employing bacteria can be used to produce biogas. Biogas is used as an energy source. Alternatively, food wastes can be valorized for the production of energy by using different common techniques such as composting, recycling and incineration. Although, these processes are capable of converting food waste into fuels and value-added products development of greener and advanced technologies are required [3].
Biofuel production is rapidly growing as the world encounters pollution problems due to burning of petroleum and coal based fuels. In addition, petroleum fuels are finite reserves and most of the petroleum reserves are geographically located in politically unstable countries. This reinforces the fact that alternative fuels are important from both environmental and energy security point of view. Along this line, many countries are formulating energy policies for the production of renewable energy.
At present, biofuels such as biodiesel and bioethanol are largely produced from edible food materials [4]. Various edible plant oils from soybean, rapeseed and canola oils are used for the preparation of biodiesel. Whereas, ethanol can be produced from a variety of feedstocks such as sugar cane, bagasse, sugar beet, grain, switchgrass, barley, potatoes, molasses, corn, stover, wheat and many other sources rich in carbohydrate [4],[5]. Chemical biodiesel production process is called transesterification [4],[5]. During transesterification, the tri-,di-and mono-glycerides react with methanol in the presence of a catalyst to produce biodiesel. On the other hand, the production of bioethanol process involves pretreatment, enzymatic hydrolysis, fermentation and distillation steps. Preparation of biofuels from edible food materials is attributed for the reason of food scarcity and a food vs fuel debate is already raging [6]. Alternatively, nonedible feedstocks can be used for the production of biofuels. Jatropha, Pongamia and other nonedible plant oils are already used for the preparation of biodiesel [5]. Along this line, nonedible lignocellulosic biomass is also employed for the production of bioethanol [7].
Food waste is a well-known nonedible source of lipids, carbohydrates, amino acids and phosphates [8],[9]. Reserach in our laboratory reveals that bakery and mixed food wastes contain significant amount of lipids and carbohydrates [8],[9]. Depending on the source of food waste the average lipid content was around 30% and the average carbohydrate content was around 50% [8],[9]. Different types of food wastes can be hydrolysed enzymatically to produce food hydrolysate and lipids [8],[9]. The food hydrolysate was rich in carbohydrate and can be used for the production of bioethanol; whereas, the obtained lipid can be converted to biodiesel (Figure 1).
thumbnailFigure 1. Recycling of food waste into biodiesel and bioethanol.
Along this line, noodle is a common starch based food material. In South Korea around 3 billion packages of instant noodle were consumed in 2011 and more than 2,100 tons of instant noodle residues were disposed as waste. Kim et al. have used instant noodle waste for the production of biofuels [10],[11]. Kim et al. recovered the oil from noodle waste by extraction using nonpolar hexane as a solvent. From 100 g of noodle waste 83 g of purified starch and 5 g of oil was obtained. The obtained oil free starch residue was used for simultaneous saccharification and fermentation process for the production of bioethanol [10],[11]. The hexane extract was evaporated and the obtained oil was reacted with methanol in the presence of acid and alkali catalysts for the preparation of biodiesel. One major limitation of this process is that the excess use of hexane during the extraction of oil from food waste [10],[11]. The Centers for Disease Control classifies n-hexane as a neurotoxin. It is also listed as a “hazardous air pollutant” by the Environmental Protection Agency as it helps in the formation of ozone at the ground level which is primarily responsible for smog. Nevertheless, these experiments demonstrate the potential utilization of waste noodle waste as a resource for biofuel production.
Research on the use of food wastes for the production of biofuel is becoming attractive in different countries. Sulaiman et al. have conceptualized a halal biorefinery for the production of fuels and value added products in Malaysia [12]. Yao et al. from the Chinese Academy of Sciences investigated the application of food waste to generate hydrolysates for the production of bioethanol [13]. In this context, potato peel is a well-known waste generated by the potato industries in Europe [14]. As reported by Christakopoulos.et al. this “no value” potato peel waste was converted to bioethanol using environmentally benign biocatalytic methods [13]. Researchers have also used household food waste for the production of bioethanol. During this process liquefaction and saccharification methods were employed to increase both ethanol production and productivity of the process [15]. Subsequently, fermentation of the remaining solids obtained from this process was performed to increase the overall yield of ethanol [15].
It is clear from the above discussion that environmental pollution and upcoming shortage of fossil fuels have turned the attention of researchers largely on the utilization of renewable feedstocks. Additionally, scientists and policy makers are devoting much efforts to use nonedible and zero cost food wastes for fuel and energy production to reduce the direct competition between fuel and food. It is known that food wastes are generated in large quantities and their handling is a challenge. As discussed earlier, these food wastes are potential resources as they contain substantial amount of carbohydrates and lipids. Thus, zero value food waste can be used as a resource for the production of low-cost biofuels. Research by various groups is currently underway for the production of biodiesel and bioethanol from food waste [16]-[22]. So far, the “proof of concept” for the synthesis and characterization of biofuels from different food wastes has been established [16]-[22].
Currently, technologies are available for the production of biodiesel and bioethanol in industrial scale. In this regard, a pilot scale production of ethanol from food waste using Saccharomyces cerevisiae H058 is already reported [22]. Nevertheless, low cost, greener and advanced technologies are needed for the production of fuels from food wastes [3],[23]. The industrial production of biofuel from food waste is largely depended on i) availability of food waste, ii) efficiency of hydrolyis process, iii) the amount of lipid and carbohydrate obtained from food waste, and iv) efficiency of fermentation and transesterification methods. In Hong Kong, Taiwan, Korea, US and in many other European countries plenty of food wastes are available. Thus, the future work should be primarily focused on large scale pretreatment and hydrolysis of food waste for production of lipid and sugar enriched hydrolysate. Several microorganisms and enzymes are known to hydrolyze food waste to carbohydrate, lipid, amino acids and phosphates. The catalytic efficiencies of the existing biocatalysts can be tested for the large scale hydrolysis of food wastes. Afterwards, the commercially available technologies can be employed for the production of biodiesel and bioethanol.
To make biofuels economically viable, the business and scientific communities and policy makers should come together to start a joint venture into the business of converting food waste into fuels. With proper financial and policy based supports from government, food waste biorefineries can be realized. In this context, it is particularly important to overcome the existing technological challenges of conventional food waste valorization methods. Simultaneously, it is crucial to develop environmental friendly and cost effective recycling methods that can convert food wastes into biofuels and chemicals [24]-[26]. Recently, combi-protein coated microcrystals of lipases are used for the production of biodiesel from oil of spent coffee grounds [27]. In this regard, both chemo- and biocatalytic methods can be explored for the preparation of biofuels from food waste [27]-[31].
The electronic version of this article is the complete one and can be found online at: http://www.sustainablechemicalprocesses.com/content/2/1/22
image source: 10 Green vehicles that run on food waste
Understanding acute kidney injury in low resource settings: a step forward
* Corresponding author: Fredric O Finkelstein fof@comcast.net
Author Affiliations
Yale University, 136 Sherman Avenue, New Haven, USA
BMC Nephrology 2015, 16:5 doi:10.1186/1471-2369-16-5
<More about author>
The International Society of Nephrology has recently set a goal of eliminating preventable or treatable deaths from acute kidney injury (AKI) by 2025—the “0X25” initiative. Programmatic implementation in low resource settings (LRS) is a key mission of this initiative. But a major challenge in designing effective programs to treat AKI in LRS is that we do not know the extent and nature of the problem.
Main text
Few studies describe the epidemiology of AKI in LRS [1-3]. In a 2013 meta-analysis examining 154 studies on the incidence of AKI, the authors could identify only two adult studies from LRS which used a standard definition of AKI with sample sizes exceeding 500 adults or 50 children [1]. A number of factors contribute to this paucity of good quality data. In LRS, late presentation of patients to tertiary care centers is common. Often, limited resources force a difficult choice between spending money on serial laboratory tests or treatment, resulting in ascertainment bias. Even when an attempt is made to study AKI systematically, the limitations on laboratory services with rapid turn-around sometimes lead to use of alternate AKI definitions, since current consensus AKI definitions are partly based on serial creatinine measurements. Furthermore, most centers still use paper rather than electronic medical records, making data extraction less efficient and multicenter collaboration more difficult. As a result, many studies are often single center with small sample sizes [2]. They often do not make it to publication in high profile journals, but instead are published in regional journals which are not readily accessible to the general medical community.
Thus, the study published in this issue of BMC Nephrology by Bagasha et al. begins to fill an important data gap [4]. Using the Acute Kidney Injury Network definition, the study describes the epidemiology and correlates of AKI among 387 adult patients admitted to Mulago National Referral Hospital (the largest hospital in Uganda) who fit criteria for a diagnosis of sepsis. The prevalence of AKI as assessed at a single time point was 16%, with 46% of patients developing severe AKI. Overall mortality was 21%.
The numbers presented in this study, while stark, are likely underestimates for several reasons. The authors excluded patients with chronic kidney disease, a group at high-risk for AKI in the setting of sepsis. The assessment for AKI occurred only at enrollment, and patients who may have developed AKI in their clinical course were not included. Moving away from a tertiary care center, we can speculate that in rural or other urban centers without academic support, the diagnosis of AKI is often missed or delayed, and outcomes will likely be even worse.
Nonetheless, the study by Bagasha et al. provides an important snapshot of the patients admitted for sepsis likely to develop AKI in Uganda: most are young, have HIV, and may well have used herbal medications prior to admission. We can extrapolate that such a patient will likely not receive intensive unit care, and if he/she develops severe AKI, the likelihood of death will approach 40%. Dialysis initiation will occur rarely or not at all, even at the largest referral center in Uganda.
This picture contrasts with the one seen in high-resource settings, where AKI is most often hospital acquired, in patients with multiorgan failure and/or multiple co-morbidities who have been exposed to polypharmacy and invasive procedures. In a study of 22 centers in the U.S., Canada, and Saudi Arabia, patients with AKI were often critically-ill, with close to 40% having 2 or more co-morbidities [5]. Average age of patients with sepsis-related AKI exceeded 60 years in this and other reports from Italy [6], Germany [7], and New Zealand [8]. Dialytic support was used in 2-10% of AKI patients [1,6-8].
The differences between AKI in LRS, as noted in the Bagasha et al. study and other single center studies, and high resource settings are important to note for several reasons. First, AKI is potentially more “preventable and treatable”. Appropriate hygiene may decrease the incidence and spread of diarrheal diseases. Volume depletion, when recognized sufficiently early, could potentially be addressed by inexpensive oral rehydration strategies in the community or intravenous fluids in the hospital. Given that the majority of patients in the study by Bagasha et al. had received <1L of fluid at the time of diagnosis, how many episodes of AKI could have been prevented if appropriate hemodynamic support were provided? Early administration of antibiotics or antimalarials could also help address potentially remediable diseases. Avoidance of nephrotoxins, another cornerstone of the management of critically ill patients, can also make a major difference in limiting the incidence and/or progression of AKI, especially given that close to 25% of patients with AKI used herbal medications prior to their admission to the hospital.
Second, the patients are young. Nearly half of patients experiencing AKI in the Bagasha et al. study were younger than 40 years; a similarly young age distribution has been reported from South Africa [9] and Sri Lanka [10]. Few patients have co-morbidities other than HIV. Often the kidney is the only failed organ. These factors contribute to a higher likelihood of full renal and overall recovery, and return of young individuals to become active members of society.
Third, if patients do progress to AKI in LRS, many have no options for dialysis therapy. Dialytic support is limited by the lack of resources, lack of available equipment, and lack of funding necessary to support such therapy. Patients are often asked to pay the high costs of treatment—costs which are generally beyond their means. These sobering facts make the prevention or at least the mitigation of AKI even more important.
Thus, the 0X25 initiative has been embraced by the nephrology community worldwide, and strategies are actively being organized to address the problem of AKI in LRS. The cornerstones of the program involve identifying the causes of AKI, understanding what resources are available to treat patients at risk for AKI, developing education programs to increase the awareness of the importance of early diagnosis and treatment of AKI, and then developing appropriate treatment strategies.
Understanding how to approach education and treatment strategies requires that plans be thought of in the structural, cultural, and financial context of the setting in which these programs are being developed. Thus, in a thoughtful and detailed review of the Millenium Village Project, Nina Munk provides a cautionary story [11]. Munk, who spent 6 years examining the impact of the Millenium Project in Africa, describes the dangers inherent in imposing outside theories on the complex and ever-changing lives of African villagers. She reflects on the cultural difficulties that outside aid efforts are likely to encounter in rural Africa. In terms of renal replacement therapy, the use of peritoneal dialysis (PD) is now being again recognized as an acceptable, cost-efficient, and technologically simple therapy to be used not only in LRS but in high income countries as well [12,13]. Thus, Chionh et al. have suggested that outcomes with PD for the treatment of AKI are as good as with extracorporeal therapies [13]. Comprehensive guidelines for the use of PD to treat patients (children and adults) with AKI have recently been published [13]. The feasibility of using PD to treat patients with AKI in LRS is now well documented with excellent outcomes being reported in recent publications from Tanzania [14], Sudan [15], and Nigeria [16]. The Saving Young Lives Program has been actively helping support and develop PD programs to treat patients with AKI in LRS [17].
The electronic version of this article is the complete one and can be found online at: http://www.biomedcentral.com/1471-2369/16/5
image source: Kent Surrey Sussex Academic Health Science Network
Toward Scalable Systems for Big Data Analytics: A Technology Tutorial
Do you know what is big data? the following article published on IEEE will help you more about the history and related technology behind this new term.
"Recent technological advancements have led to a deluge of data from distinctive domains (e.g., health care and scientific sensors, user-generated data, Internet and financial companies, and supply chain systems) over the past two decades. The term big data was coined to capture the meaning of this emerging trend. In addition to its sheer volume, big data also exhibits other unique characteristics as compared with traditional data. For instance, big data is commonly unstructured and require more real-time analysis. This development calls for new system architectures for data acquisition, transmission, storage, and large-scale data processing mechanisms."
Corresponding author: Yonggang Wen ygwen@ntu.edu.sg
DOI 10.1109/ACCESS.2014.2332453
You can find the complete article (open access) below:
The science of taste
Correspondence: Ole G Mouritsen ogm@memphys.sdu.dk
MEMPHYS, Center for Biomembrane Physics and TASTEforLIFE, Department of Physics, Chemistry, and Pharmacy, University of Southern Denmark, Campusvej 55, Odense M, DK-5230, Denmark
Flavour 2015, 4:18 doi:10.1186/s13411-014-0028-3
[More about author]
Homepage at University of Southern Denmark
Editorial
In contrast to smell and the olfactory system, for which the 2004 Nobel Prize in Physiology and Medicine was awarded to Richard Axel and Linda Buck for their discovery of odorant receptors and the organization of the olfactory system [1], our knowledge of the physiological basis for the taste system is considerably less developed [2]. Some progress has been obtained over the last decade by the finding of receptors or receptor candidates for all five basic tastes, bitter, sweet, umami, sour, and salty. The receptors for bitter, sweet, and umami appear to belong to the same superfamily of G-protein-coupled receptors, whereas the receptor for salty is an ion channel. The receptor function for sour is the least understood but may involve some kind of proton sensing.
Notwithstanding the prominent status of physiology of taste and its molecular underpinnings, the multisensory processing and integration of taste with other sensory inputs (sight, smell, sound, mouthfeel, etc.) in the brain and neural system have also received an increasing attention, and an understanding is emerging of how taste relates to learning, perception, emotion, and memory [3]. Similarly, the psychology of taste and how taste dictates food choice, acceptance, and hedonic behavior are in the process of being uncovered [4]. Development of taste preferences in children and gustatory impairment in sick and elderly are now studied extensively to understand the nature of taste and the use of this insight to improve the quality of life.
Finally, a new direction has manifested itself in recent years where scientists and creative chefs apply scientific methods to gastronomy in order to explore taste in traditional and novel dishes and use physical sciences to characterize foodstuff, cooking, and flavor [5]-[8].
Noting that in general our understanding of taste is inferior to our knowledge of the other human senses, an interdisciplinary symposium, The Science of Taste, took place in August 2014 and brought together an international group of scientists and practitioners from a range of different disciplines (biophysics, physiology, sensory sciences, neuroscience, nutrition, psychology, epidemiology, food science, gastronomy, gastroscience, and anthropology) to discuss progress in the science of taste. As a special feature, the symposium organized two tasting events arranged by leading chefs, demonstrating the interaction between creative chefs and scientists.
The symposium led to the following special collection of papers accounting for our current knowledge about the science of taste. The collection includes a selection of opinion articles, short reports, and reviews, in addition to three research papers.
The papers deal with the following topics: the comparative biology of taste [9]; fat as a basic taste [10]; umami taste in relation to gastronomy [11]; the mechanism of kokumi taste [12]; geography as a starting point for deliciousness [13], temporal design of taste and flavor [14]; the pleasure principle of flavors [15]; taste as a cultural activity [16]; taste preferences in primary school children [17]; taste and appetite [18]; umami taste in relation to health [19]; taste receptors in the gastrointestinal tract [20]; neuroenology and the taste of wine [21]; the brain mechanisms behind pleasure [22]; the importance of sound for taste [23]; as well the effect of kokumi substances on the flavor of particular food items [24],[25].
The electronic version of this article is the complete one and can be found online at: http://www.flavourjournal.com/content/4/1/18
The whole collection of article can be found here:
http://www.flavourjournal.com/series/the_science_of_taste
image source: Tip of the Tongue: Humans May Taste at Least 6 Flavors
Extending reference assembly models
Corresponding authors:
Deanna M Church deanna.church@personalis.com
Valerie A Schneider schneiva@ncbi.nlm.nih.gov
Richard Durbin rd@sanger.ac.uk
Paul Flicek flicek@ebi.ac.uk
Genome Biology 2015, 16:13 doi:10.1186/s13059-015-0587-3
Background
One of the flagship products of the Human Genome Project (HGP) was a high-quality human reference assembly [1]. This assembly, coupled with advances in low-cost, high-throughput sequencing, has allowed us to address previously inaccessible questions about population diversity, genome structure, gene expression and regulation [2]-[5]. It has become clear, however, that the original models used to represent the reference assembly inadequately represent our current understanding of genome architecture.
The first assembly models were designed for simple ‘linear’ genome sequences, with little sequence variation and even less structural diversity. The design fit the understanding of human variation at the time the HGP began [6]. The HGP constructed the reference assembly by collapsing sequences from over 50 individuals into a single consensus haplotype representation of each chromosome. Employing a clone-based approach, the sequence of each clone represented a single haplotype from a given donor. At clone boundaries, however, haplotypes could switch abruptly, creating a mosaic structure. This design introduced errors within regions of complex structural variation, when sequences unique to one haplotype prevented construction of clone overlaps. The assembly therefore inadvertently included multiple haplotypes in series in some regions [7]-[9].
The Genome Reference Consortium (GRC) began stewardship of the reference assembly in 2007. The GRC proposed a new assembly model that formalized the inclusion of ‘alternative sequence paths’ in regions with complex structural variation, and then released GRCh37 using this new model [10]. The release of GRCh37 also marked the deposition of the human reference assembly to an International Nucleotide Sequence Database Collaboration (INSDC) database, providing stable, trackable sequence identifiers, in the form of accession and version numbers, for all sequences in the assembly. The GRC developed an assembly model that was incorporated into the National Centre for Biotechnology Information (NCBI) and European Nucleotide Archive (ENA) assembly database that provides a stable identifier for the collection of sequences and the relationship between these sequences that comprise an assembly [11]. Subsequent minor assembly releases added a number of ‘fix patches’ that could be used to resolve mistakes in the reference sequence, as well as ‘novel patches’ that are new alternative sequence representations [10].
The new assembly model presents significant advances to the genomics community, but, to realize those advances, we must address many technical challenges. The new assembly model is neither haploid nor diploid - instead, it includes additional scaffold sequences, aligned to the chromosome assembly, that provide alternative sequence representations for regions of excess diversity. Widely used alignment programs, variant discovery and analysis tools, as well as most reporting formats, expect reads and features to have a single location in the reference assembly as they were developed using a haploid assembly model. Many alignment and analysis tools penalize reads that align to more than one location under the assumption that the location of these reads cannot be resolved owing to paralogous sequences in the genome. These tools do not distinguish allelic duplication, added by the alternative loci, from paralogous duplication found in the genome, thus confounding repeat and mappability calculations, paired-end placements and downstream interpretation of alignments in regions with alternative loci.
To determine the efforts needed to facilitate use of the full assembly, the GRC organized a workshop in conjunction with the 2014 Genome Informatics meeting in Cambridge, UK (http://www.slideshare.net/GenomeRef webcite). Participants identified challenges presented by the new assembly model and discussed ways forward that we describe here.
Towards the graph of human variation
A graph structure is a natural way to represent a population-based genome assembly, with branches in the graph representing all variation found within the source sequences. Most assembly programs internally use a graph representation to build the assembly, but ultimately produce a flattened structure for use by downstream tools [12]-[14]. Recently, formal proposals for representing a population-based reference graph have been described [15]-[17]. The newly formed Global Alliance for Genomics and Health (GA4GH) is leading an effort to formalize data structures for graph-based reference assemblies, but it will likely take years to develop the infrastructure and analysis tools needed to support these new structures and see their widespread adoption across the biological and clinical research communities [18].
The introduction of alternative loci into the assembly model provides a stepping-stone towards a full graph-based representation of a population-based reference genome. The alternative loci provided by the GRC are based on high-quality, finished sequence. Although it is not feasible to represent all known variation using the alternative locus scheme, this model does allow us to better represent regions with extreme levels of diversity. Alternative loci are not meant to represent all variation within a population, but rather provide an immediate solution for adding sequences missing from the chromosome assembly. In practice, alternative locus addition is limited by the availability of high-quality genomic sequence, and the GRC has focused on representing sequence at the most diverse regions, such as the major histocompatibility complex (MHC). The representation of all population variation is better suited to a graph-based representation. The high quality of the sequence at these locations provides robust data to test graph implementations. Additionally, because both NCBI and Ensembl have annotated these sequences, we can also begin to address how to annotate graph structures at these complex loci.
While GRCh37 had only three regions containing nine alternative locus sequences, GRCh38 has 178 regions containing 261 alternative locus sequences, collectively representing 3.6 Mbp of novel sequence and over 150 genes not represented in the primary assembly (Table 1). The increased level of alternative sequence representation intensifies the urgency to develop new analysis methods to support inclusion of these sequences. Inclusion of all sequences in the reference assembly allows us to better analyze these regions with potentially modest updates to currently used tools and reporting structures. Although the addition of the alternative loci to current analysis pipelines might lead to only modest gains in analysis power on a genome-wide scale, some loci will see considerable improvement owing to the addition of significant amounts of sequence that cannot be represented accurately in the chromosome assembly (Figure 1).
Omission of the novel sequence contained in the alternative loci can lead to off-target sequence alignments, and thus incorrect variant calls or other errors, when a sample containing the alternative allele is sequenced and aligned to only the primary assembly. Using reads simulated from the unique portion of the alternative loci, we found that approximately 75% of the reads had an off-target alignment when aligned to the primary assembly alone. This finding was consistent using different alignment methods [10]. The 1000 Genomes Project also observed the detrimental effect of missing sequences and developed a ‘decoy’ sequence dataset in an effort to minimize off-target alignments [19],[20]. Much of this decoy has now been incorporated into GRCh38, and analysis of reads taken from 1000 Genomes samples that previously mapped only to the decoy shows that approximately 70% of these now align to the full GRCh38, with approximately 1% of these reads aligning only to the alternative loci (Figure 2).
We foresee many computational approaches that allow the inclusion of all assembly sequences in analysis pipelines. To better support exploration in this area, we propose some improvements to standard practices and data structures that will facilitate future development.
Enhancement of standard reporting formats (such as BAM/CRAM, VCF/BCF, GFF3) so that they can accommodate features with multiple locations. Doing so while maintaining the allelic relationship between these features is crucial [21]-[24].
Adoption of standard sequence identifiers for sequence analysis and reporting. Using shorthand identifiers (for example, ‘chr1’ or ‘1’) to indicate the sequence is imprecise and also ignores the presence of other sequences in the assembly. In many cases, other top-level sequences, such as unlocalized scaffolds, patches and alternative loci, have a chromosome assignment but not chromosome coordinates. These sequences are independent of the chromosome assembly coordinate system and have their own coordinate space. Alternative loci are related to the chromosome coordinates through alignment to the chromosome assembly. Developing a structure that treats all top-level sequences as first-class citizens during analysis is an important step towards adopting use of the full assembly in analysis pipelines.
Curation of multiple sequence alignments of the alternative loci to each other and the primary path. Currently, pairwise alignments of the alternative loci to the chromosome assembly are available to provide the allelic relationship between the alternative locus and the chromosome. However, these pairwise alignments do not allow for the comparison of alternative loci in a given region to each other. These alignments can also be used to develop graph structures. The relationship of the allelic sequences within a region helps define the assembly structure, and the community should work from a single set of alignments. These should be distributed with the GRC assembly releases.
Recently, the GRC has released a track hub [25] that allows for the distribution of GRC data using standard track names and content (http://ngs.sanger.ac.uk/production/grit/track_hub/hub.txt webcite). Additionally, the GRC has created a GitHub page to track development of tools and resources that facilitate use of the full assembly (https://github.com/GenomeRef/SoftwareDevTracking webcite).
Concluding remarks
As we gain understanding of biological systems, we must update the models we use to represent these data. This can be difficult when the model supports common infrastructure and analysis tools used by a large swath of the scientific community. However, this growth is crucial in order to move the scientific community forward. While adoption of this new model will take substantial effort, doing so is an important step for the human genetics and broader genomics communities. We now have an opportunity and imperative to revisit old assumptions and conventions to develop a more robust analysis framework. The use of all sequences included in the reference will allow for improved genomic analyses and understanding of genomic architecture. Additionally, this new assembly model allows us to take a small step towards the realization of a graph-based assembly representation. The evolution of the assembly model allows us to improve our understanding of genomic architecture and provides a framework for boosting our understanding of how this architecture impacts human development and disease.
The electronic version of this article is the complete one and can be found online at: http://genomebiology.com/2015/16/1/13
image source: http://www.ncbi.nlm.nih.gov/projects/genome/assembly/grc/human/
Deanna M Church deanna.church@personalis.com
Valerie A Schneider schneiva@ncbi.nlm.nih.gov
Richard Durbin rd@sanger.ac.uk
Paul Flicek flicek@ebi.ac.uk
Genome Biology 2015, 16:13 doi:10.1186/s13059-015-0587-3
Background
One of the flagship products of the Human Genome Project (HGP) was a high-quality human reference assembly [1]. This assembly, coupled with advances in low-cost, high-throughput sequencing, has allowed us to address previously inaccessible questions about population diversity, genome structure, gene expression and regulation [2]-[5]. It has become clear, however, that the original models used to represent the reference assembly inadequately represent our current understanding of genome architecture.
The first assembly models were designed for simple ‘linear’ genome sequences, with little sequence variation and even less structural diversity. The design fit the understanding of human variation at the time the HGP began [6]. The HGP constructed the reference assembly by collapsing sequences from over 50 individuals into a single consensus haplotype representation of each chromosome. Employing a clone-based approach, the sequence of each clone represented a single haplotype from a given donor. At clone boundaries, however, haplotypes could switch abruptly, creating a mosaic structure. This design introduced errors within regions of complex structural variation, when sequences unique to one haplotype prevented construction of clone overlaps. The assembly therefore inadvertently included multiple haplotypes in series in some regions [7]-[9].
The Genome Reference Consortium (GRC) began stewardship of the reference assembly in 2007. The GRC proposed a new assembly model that formalized the inclusion of ‘alternative sequence paths’ in regions with complex structural variation, and then released GRCh37 using this new model [10]. The release of GRCh37 also marked the deposition of the human reference assembly to an International Nucleotide Sequence Database Collaboration (INSDC) database, providing stable, trackable sequence identifiers, in the form of accession and version numbers, for all sequences in the assembly. The GRC developed an assembly model that was incorporated into the National Centre for Biotechnology Information (NCBI) and European Nucleotide Archive (ENA) assembly database that provides a stable identifier for the collection of sequences and the relationship between these sequences that comprise an assembly [11]. Subsequent minor assembly releases added a number of ‘fix patches’ that could be used to resolve mistakes in the reference sequence, as well as ‘novel patches’ that are new alternative sequence representations [10].
The new assembly model presents significant advances to the genomics community, but, to realize those advances, we must address many technical challenges. The new assembly model is neither haploid nor diploid - instead, it includes additional scaffold sequences, aligned to the chromosome assembly, that provide alternative sequence representations for regions of excess diversity. Widely used alignment programs, variant discovery and analysis tools, as well as most reporting formats, expect reads and features to have a single location in the reference assembly as they were developed using a haploid assembly model. Many alignment and analysis tools penalize reads that align to more than one location under the assumption that the location of these reads cannot be resolved owing to paralogous sequences in the genome. These tools do not distinguish allelic duplication, added by the alternative loci, from paralogous duplication found in the genome, thus confounding repeat and mappability calculations, paired-end placements and downstream interpretation of alignments in regions with alternative loci.
To determine the efforts needed to facilitate use of the full assembly, the GRC organized a workshop in conjunction with the 2014 Genome Informatics meeting in Cambridge, UK (http://www.slideshare.net/GenomeRef webcite). Participants identified challenges presented by the new assembly model and discussed ways forward that we describe here.
Towards the graph of human variation
A graph structure is a natural way to represent a population-based genome assembly, with branches in the graph representing all variation found within the source sequences. Most assembly programs internally use a graph representation to build the assembly, but ultimately produce a flattened structure for use by downstream tools [12]-[14]. Recently, formal proposals for representing a population-based reference graph have been described [15]-[17]. The newly formed Global Alliance for Genomics and Health (GA4GH) is leading an effort to formalize data structures for graph-based reference assemblies, but it will likely take years to develop the infrastructure and analysis tools needed to support these new structures and see their widespread adoption across the biological and clinical research communities [18].
The introduction of alternative loci into the assembly model provides a stepping-stone towards a full graph-based representation of a population-based reference genome. The alternative loci provided by the GRC are based on high-quality, finished sequence. Although it is not feasible to represent all known variation using the alternative locus scheme, this model does allow us to better represent regions with extreme levels of diversity. Alternative loci are not meant to represent all variation within a population, but rather provide an immediate solution for adding sequences missing from the chromosome assembly. In practice, alternative locus addition is limited by the availability of high-quality genomic sequence, and the GRC has focused on representing sequence at the most diverse regions, such as the major histocompatibility complex (MHC). The representation of all population variation is better suited to a graph-based representation. The high quality of the sequence at these locations provides robust data to test graph implementations. Additionally, because both NCBI and Ensembl have annotated these sequences, we can also begin to address how to annotate graph structures at these complex loci.
While GRCh37 had only three regions containing nine alternative locus sequences, GRCh38 has 178 regions containing 261 alternative locus sequences, collectively representing 3.6 Mbp of novel sequence and over 150 genes not represented in the primary assembly (Table 1). The increased level of alternative sequence representation intensifies the urgency to develop new analysis methods to support inclusion of these sequences. Inclusion of all sequences in the reference assembly allows us to better analyze these regions with potentially modest updates to currently used tools and reporting structures. Although the addition of the alternative loci to current analysis pipelines might lead to only modest gains in analysis power on a genome-wide scale, some loci will see considerable improvement owing to the addition of significant amounts of sequence that cannot be represented accurately in the chromosome assembly (Figure 1).
Omission of the novel sequence contained in the alternative loci can lead to off-target sequence alignments, and thus incorrect variant calls or other errors, when a sample containing the alternative allele is sequenced and aligned to only the primary assembly. Using reads simulated from the unique portion of the alternative loci, we found that approximately 75% of the reads had an off-target alignment when aligned to the primary assembly alone. This finding was consistent using different alignment methods [10]. The 1000 Genomes Project also observed the detrimental effect of missing sequences and developed a ‘decoy’ sequence dataset in an effort to minimize off-target alignments [19],[20]. Much of this decoy has now been incorporated into GRCh38, and analysis of reads taken from 1000 Genomes samples that previously mapped only to the decoy shows that approximately 70% of these now align to the full GRCh38, with approximately 1% of these reads aligning only to the alternative loci (Figure 2).
We foresee many computational approaches that allow the inclusion of all assembly sequences in analysis pipelines. To better support exploration in this area, we propose some improvements to standard practices and data structures that will facilitate future development.
Enhancement of standard reporting formats (such as BAM/CRAM, VCF/BCF, GFF3) so that they can accommodate features with multiple locations. Doing so while maintaining the allelic relationship between these features is crucial [21]-[24].
Adoption of standard sequence identifiers for sequence analysis and reporting. Using shorthand identifiers (for example, ‘chr1’ or ‘1’) to indicate the sequence is imprecise and also ignores the presence of other sequences in the assembly. In many cases, other top-level sequences, such as unlocalized scaffolds, patches and alternative loci, have a chromosome assignment but not chromosome coordinates. These sequences are independent of the chromosome assembly coordinate system and have their own coordinate space. Alternative loci are related to the chromosome coordinates through alignment to the chromosome assembly. Developing a structure that treats all top-level sequences as first-class citizens during analysis is an important step towards adopting use of the full assembly in analysis pipelines.
Curation of multiple sequence alignments of the alternative loci to each other and the primary path. Currently, pairwise alignments of the alternative loci to the chromosome assembly are available to provide the allelic relationship between the alternative locus and the chromosome. However, these pairwise alignments do not allow for the comparison of alternative loci in a given region to each other. These alignments can also be used to develop graph structures. The relationship of the allelic sequences within a region helps define the assembly structure, and the community should work from a single set of alignments. These should be distributed with the GRC assembly releases.
Recently, the GRC has released a track hub [25] that allows for the distribution of GRC data using standard track names and content (http://ngs.sanger.ac.uk/production/grit/track_hub/hub.txt webcite). Additionally, the GRC has created a GitHub page to track development of tools and resources that facilitate use of the full assembly (https://github.com/GenomeRef/SoftwareDevTracking webcite).
Concluding remarks
As we gain understanding of biological systems, we must update the models we use to represent these data. This can be difficult when the model supports common infrastructure and analysis tools used by a large swath of the scientific community. However, this growth is crucial in order to move the scientific community forward. While adoption of this new model will take substantial effort, doing so is an important step for the human genetics and broader genomics communities. We now have an opportunity and imperative to revisit old assumptions and conventions to develop a more robust analysis framework. The use of all sequences included in the reference will allow for improved genomic analyses and understanding of genomic architecture. Additionally, this new assembly model allows us to take a small step towards the realization of a graph-based assembly representation. The evolution of the assembly model allows us to improve our understanding of genomic architecture and provides a framework for boosting our understanding of how this architecture impacts human development and disease.
The electronic version of this article is the complete one and can be found online at: http://genomebiology.com/2015/16/1/13
image source: http://www.ncbi.nlm.nih.gov/projects/genome/assembly/grc/human/










