Topics (203k) Articles (205k) Users (300) Discussions (237) Comments (383) Files (715) New article
Association mapping, also known as linkage disequilibrium mapping, is a genetic analysis method used to identify the relationship between genetic markers and traits of interest in a population. It is particularly useful in understanding the genetic basis of complex traits, such as those influenced by multiple genes and environmental factors. ### Key Concepts: 1. **Genetic Markers**: These are specific sequences in the genome, such as single nucleotide polymorphisms (SNPs), that vary among individuals.
Allelotype is a genetic concept referring to the specific pattern of alleles (variant forms of a gene) present in an individual's genome, especially concerning the variation in alleles that are associated with certain traits or diseases. The term is often used in the context of genetic studies to analyze the distribution and inheritance of alleles among populations, and it can help in identifying genetic predispositions to certain conditions.
Additive disequilibrium and the Z statistic are concepts used in population genetics and evolutionary biology, particularly in the study of genetic variation and allele frequency distributions. ### Additive Disequilibrium: Additive disequilibrium refers to the deviation from expected allele frequencies in a population, often observed when there are non-random associations between alleles at different genetic loci. This can be a result of various evolutionary forces such as natural selection, genetic drift, migration, or non-random mating.
Statistical geneticists are specialists who apply statistical methods and techniques to understand genetic data and contribute to the field of genetics. Their work involves analyzing data that can help to uncover the relationships between genetic variation and traits or diseases, thereby advancing our understanding of the genetic basis of various biological processes.
Quantitative genetics is a branch of genetics that deals with the inheritance of traits that are determined by multiple genes (polygenic traits) rather than a single gene. This field focuses on understanding how genetic and environmental factors contribute to the variation in traits within a population. Key aspects of quantitative genetics include: 1. **Traits**: Quantitative traits are typically measurable and can include characteristics such as height, weight, yield in crops, or susceptibility to diseases.
The Yamartino method is a well-known approach used for estimating the parameters of statistical models, particularly in the field of time series analysis. It focuses on time series data where the observations are influenced by seasonality or periodic effects. The method involves decomposing the time series into its components—trend, seasonality, and error. One of the main applications of the Yamartino method is in forecasting, where it helps in providing more accurate predictions by taking into account the seasonal structure of the data.
Repeated median regression is a robust statistical method used for estimating the central tendency of a set of data points, specifically when dealing with repeated measures or grouped data. The method is particularly useful in situations where the data may contain outliers or do not meet the assumptions of traditional regression techniques, such as normality. In repeated median regression, the median is computed for each group of repeated measures rather than the mean, which makes this approach less sensitive to extreme values.
Random Sample Consensus (RANSAC) is an iterative algorithm used in robust estimation to fit a mathematical model to a set of observed data points. It is particularly useful when dealing with data that may contain a significant proportion of outliers—data points that do not conform to the expected model. Here’s how the RANSAC algorithm generally works: 1. **Random Selection**: Randomly select a subset of the original data points.
The Pseudo-marginal Metropolis-Hastings (PMMH) algorithm is a Markov Chain Monte Carlo (MCMC) method used for sampling from complex posterior distributions, particularly in Bayesian inference settings. It is especially useful when the likelihood function is intractable or computationally expensive to evaluate directly. ### Overview In standard MCMC methods, a proposal distribution is used to explore the parameter space, and the acceptance criterion is based on the ratio of the posterior probabilities.
The Metropolis–Hastings algorithm is a Markov Chain Monte Carlo (MCMC) method used for sampling from probability distributions that are difficult to sample from directly. It is particularly useful in situations where the distribution is defined up to a normalization constant, making it challenging to derive samples analytically.
The Lander–Green algorithm is a method used for generating random samples from the uniform distribution over specific combinatorial objects such as integer partitions or certain types of labeled structures. It is particularly well-known for its application in generating random integer partitions efficiently. The algorithm operates by combining techniques from combinatorial enumeration and probabilistic sampling. It ensures that each possible configuration has an equal chance of being selected, which is crucial for applications in statistical analysis, simulations, and other computational problems.
Kernel-independent component analysis (KICA) is an extension of independent component analysis (ICA) that utilizes kernel methods to allow for the separation of non-linear components from data. While standard ICA is designed to separate independent sources in a linear fashion, KICA broadens this capability by applying kernel techniques, which can handle more complex relationships within the data.
Iterative Proportional Fitting (IPF), also known as Iterative Proportional Scaling (IPS) or the RAS algorithm, is a statistical method used to adjust the values in a multi-dimensional contingency table so that they meet specified marginal totals. This technique is particularly useful in fields like economics, demography, and social sciences, where researchers often work with incomplete data or need to align observed data with known populations.
HyperLogLog is a probabilistic data structure used for estimating the cardinality (the number of distinct elements) of a multiset (a collection of elements that may contain duplicates) in a space-efficient manner. It is particularly useful for applications that require approximate counts of unique items for large datasets. ### Key Features: 1. **Space Efficiency**: HyperLogLog uses significantly less memory compared to exact counting methods.
Helmert-Wolf blocking is a method used in survey geodesy and geospatial analysis for processing and adjusting measurements made on a network of points. It is named after the geodesists Friedrich Helmert and Paul Wolf, who contributed to the development of techniques for adjusting geodetic networks. In essence, Helmert-Wolf blocking is a strategy for dividing a large network of observations into smaller, more manageable segments or blocks.
Farr's laws refer to principles in epidemiology related to the relationship between health outcomes, particularly mortality rates, and the characteristics of the population being studied. Specifically, they are associated with the work of Sir Edwin Chadwick and William Farr in the 19th century, who contributed significantly to the field of public health and statistics. Farr's laws focus on the idea that the mortality rates of specific diseases can be predicted based on the age structure of a population and the spatial distribution of that population.
The False Nearest Neighbor (FNN) algorithm is a technique used primarily in the context of time series analysis and nonlinear dynamics to determine the appropriate number of embedding dimensions required for reconstructing the state space of a dynamical system. It is particularly useful in the study of chaotic systems. ### Key Concepts of the FNN Algorithm: 1. **State Space Reconstruction**: In dynamical systems, especially chaotic ones, it is often necessary to reconstruct the state space from a single-time series measurement.
The Elston–Stewart algorithm is a statistical method used for computing the likelihoods of genetic data in the context of genetic linkage analysis. It is particularly useful in the study of pedigrees, which are family trees that display the transmission of genetic traits through generations. ### Key Features of the Elston–Stewart Algorithm: 1. **Purpose**: The algorithm is designed to efficiently compute the likelihood of observing certain genotypes (genetic variants) in a family pedigree given specific genetic models.
The Count-Distinct problem is a common problem in computer science and data analysis that involves counting the number of distinct (unique) elements in a dataset. This problem often arises in database queries, data mining, and big data applications where an efficient way to determine the number of unique items is needed.
Chi-square Automatic Interaction Detection (CHAID) is a statistical technique used for segmenting a dataset into distinct groups based on the relationships between variables. It is particularly useful in exploratory data analysis, market research, and predictive modeling. CHAID is a type of decision tree methodology that utilizes the Chi-square test to determine the optimal way to split a dataset into categories.
Pinned article: Introduction to the OurBigBook Project
Welcome to the OurBigBook Project! Our goal is to create the perfect publishing platform for STEM subjects, and get university-level students to write the best free STEM tutorials ever.
Everyone is welcome to create an account and play with the site: ourbigbook.com/go/register. We belive that students themselves can write amazing tutorials, but teachers are welcome too. You can write about anything you want, it doesn't have to be STEM or even educational. Silly test content is very welcome and you won't be penalized in any way. Just keep it legal!
Intro to OurBigBook
. Source. We have two killer features:
- topics: topics group articles by different users with the same title, e.g. here is the topic for the "Fundamental Theorem of Calculus" ourbigbook.com/go/topic/fundamental-theorem-of-calculusArticles of different users are sorted by upvote within each article page. This feature is a bit like:
- a Wikipedia where each user can have their own version of each article
- a Q&A website like Stack Overflow, where multiple people can give their views on a given topic, and the best ones are sorted by upvote. Except you don't need to wait for someone to ask first, and any topic goes, no matter how narrow or broad
This feature makes it possible for readers to find better explanations of any topic created by other writers. And it allows writers to create an explanation in a place that readers might actually find it.Figure 1. Screenshot of the "Derivative" topic page. View it live at: ourbigbook.com/go/topic/derivativeVideo 2. OurBigBook Web topics demo. Source. - local editing: you can store all your personal knowledge base content locally in a plaintext markup format that can be edited locally and published either:This way you can be sure that even if OurBigBook.com were to go down one day (which we have no plans to do as it is quite cheap to host!), your content will still be perfectly readable as a static site.
- to OurBigBook.com to get awesome multi-user features like topics and likes
- as HTML files to a static website, which you can host yourself for free on many external providers like GitHub Pages, and remain in full control
Figure 2. You can publish local OurBigBook lightweight markup files to either OurBigBook.com or as a static website.Figure 3. Visual Studio Code extension installation.Figure 5. . You can also edit articles on the Web editor without installing anything locally. Video 3. Edit locally and publish demo. Source. This shows editing OurBigBook Markup and publishing it using the Visual Studio Code extension. - Infinitely deep tables of contents:
All our software is open source and hosted at: github.com/ourbigbook/ourbigbook
Further documentation can be found at: docs.ourbigbook.com
Feel free to reach our to us for any help or suggestions: docs.ourbigbook.com/#contact





