Biodiversity partitioning - Partizionamento della biodiversità

Site: LifeWatch ERIC Training Platform
Course: Introduction to partitioning and mixed models for biodiversity analysis in R
Book: Biodiversity partitioning - Partizionamento della biodiversità
Printed by: Guest user
Date: Monday, 27 July 2026, 12:13 PM

Description

Biodiversity partitioning - Partizionamento della biodiversità

Tsallis entropy and the deformed exponential transformation

  • Entropy can be thought of as the amount of uncertainty calculated from the frequency distribution of the species in a community. Popular biodiversity measures, including the number of species, Shannon and Simpson indices, can be given a rigorous framework based on the generalized Tsallis entropy [1]. While the interpretation of the Tsallis entropy is not straightforward, one can easily transform it into Hill numbers having many desirable properties and corresponding to the number of equally-frequent species that would give the same level of diversity as the data.
  • Consider a community where n individuals are sampled. Let s = 1,..., S denote the species that compose the community and ns be the number of sampled individuals of species s, with \(\sum_{s=1}^S n_s=n\).
    The probability for an individual to belong to species s is estimated by ps = ns/n. Given a discrete set of probabilities p = (p1,...,ps) and any real number q, the Tsallis entropy of order q is defined as
    $$ H_q({p}) = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_s^q \right). $$
    The number of species is the Tsallis entropy of order q = 0, while Shannon's and Simpson's indices respectively correspond to q = 1 and q = 2. The importance given to rare species decreases continuously with q.
  • Corresponding true diversity measures Dq(p), or Hill numbers, are obtained taking the deformed exponential transformation eq of the Tsallis entropy:
    $$ D_q(p)=e_q\left(H_q(p)\right) $$
    For \(H_q(p) < \frac{1}{q-1}\) the deformed exponential transformation of order \(q\) is defined as
    $$ e_q\left(H_q(p)\right) = [1+(1-q)H_q(p)]^{1 \over 1-q} $$
    with the standard exponential transformation obtained as a special case when q=1.
  • Notice that the opposite also holds, as we can define the deformed logarithm of order \(q\) as
    $$ \ln_q \left(D_q(p)\right) = \frac 1 {1-q} \left[\left(D_q(p)\right)^{1-q}-1\right]=H_q(p) $$

L'entropia di Tsallis e la trasformazione esponziale deformata

  • L'entropia puó essere pensata come la quantitá di incertezza da una distribuzione di frequenza delle specie in una comunitá. Misurazioni comuni di biodiversitá, compresi il numero delle specie, gli indici di Shannon e di Simpson, possono essere dotate di un sistema rigoroso basato sulla entropia generalizzata di Tsallis[1]. Mentre l'interpretazione dell'entropia di Tsallis non semplice, é possibile trasformarla in numeri di Hill aventi molte proprietá interessanti e la numero di specie di frequenza uguale corrispondono livelli di diversitá uguali, come nei dati.
  • Si consideri una comunitá dove n esemplari sono campionati. Siano s=1,...,S le specie che compongono la comunitá e ns sia il numero degli esemplari campionati della specie s, con \(\sum_{s=1}^S n_s=n\).
    La probabilitá per un esemplare di appartenente ad una specie s é calcolata da ps = ns/n. Dato un insieme discreto di probabilitá p = (p1,...,ps) e un numero reale q, l'entropia di Tsallis di ordine q é definita come
    $$ H_q({p}) = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_s^q \right). $$
    Il numero di specie é si ottiente quando l'entropia di Tsallis é di ordine q = 0, mentre gli indici di Shannon e di Simpson si ottengono, rispettivamente, per q = 1 e q = 2. L'importanza data ad una specie rara diminuisce continuamente con q.
  • Valori corrispondenti di vera diversitá Dq(p), o Numeri di Hill, sono ottenuti utilizzando l'equazione della trasformazione esponenziale deformata dell'entropia di Tsallis:
    $$ D_q(p)=e_q\left(H_q(p)\right) $$
    Per \(H_q(p) < \frac{1}{q-1}\) la trasformazione esponenziale deformata di ordine \(q\) é definita come
    $$ e_q\left(H_q(p)\right) = [1+(1-q)H_q(p)]^{1 \over 1-q} $$
    dove la trasformazione esponenziale standard si ottiene come caso particolare per q=1.
  • Si noti che si anche ottener l'opposto, si puó definire il logaritmo deformato di ordine \(q\) come
    $$ \ln_q \left(D_q(p)\right) = \frac 1 {1-q} \left[\left(D_q(p)\right)^{1-q}-1\right]=H_q(p) $$

Biodiversity partitioning

  • Consider a meta-community partitioned into \(I\) local communities denoted by \(i=1,\ldots,I\) Let \(n_{si}\) be the number of sampled individuals of species \(s\) in the local community \(i\), with \(\sum_{s=1}^S n_{si}=n_i\) and \(\sum_{i=1}^I n_i=n\), where \(n_i\) and \(n\) are respectively the number of sampled individuals in the \(i\)-th community and in the meta-community. Within each community \(i\), the probability for an individual to belong to species \(s\) is estimated by \(p_{si} =n_{si}/n_i\). The same probability for the meta-community is \(p_{s}\). Communities may have a weight \(w_i\), satisfying \(p_s = \sum_{i=1}^I w_ip_{si}\). The commonly-used \(w_{i}=n_{i}\) is a possible weight, but the weighting may be arbitrary (e.g. the sampled areas).

 

  • The Tsallis \(\gamma\) entropy of order \(q\) of the meta-community is defined as:
    $$ ^\gamma H_q = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_s^q \right)=\cdots=-\sum_{s=1}^S p_s^q \ln_q p_s. $$
    The Tsallis \(\alpha\) entropy of the \(i\)-th community} is:
    $$ ^\alpha _i H_q = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_{si}^q \right). $$
    The natural definition of the total \(\alpha\) entropy of the meta-community is the weighted average of local community's \(\alpha\) entropies
    $$ ^\alpha H_q = \sum_{i=1}^I w_i \;{^\alpha _i}H_q $$

 

  • \(\alpha\) and \(\gamma\) true diversity values are given by Hill numbers \(D_{q}\), i.e. the number of equally-frequent species that would give the same level of diversity as the data. According to the definition of Routledge [3] [4] Hill numbers have the following expressions for the \(\gamma\) and \(\alpha\) diversity of the meta-community:
    $$ ^\gamma D_q = \left(\sum_{s=1}^S p_s^q\right)^{\frac 1 {1-q}}=\cdots=e_q\left({^\gamma H_q}\right) $$

    $$ ^\alpha D_q = \left(\sum_{i=1}^I w_i \sum_{s=1}^S p_{si}^q\right)^{\frac 1 {q-1}} $$

 

  • Diversity measures are traditionally partitioned into \(\gamma\), \(\alpha\) and \(\beta\) and diversity [5], where
    $$ {^\gamma D_q} ={^\alpha D_q} \times {^\beta D_q} $$

 

  • Starting from the multiplicative partitioning of the \(\gamma\) diversity, the following weighted average is obtained for the generalized \(\beta\) entropy of order \(q\):
    $$ {^\beta H_q} = {^\gamma H_q}-{^\alpha H_q} = \sum_{i=1}^I w_i\;{^\beta _i H_q} $$
    where
    $$ {^\beta _i H_q}=\sum_{s=1}^S p_{si}^q \ln_q \frac{p_{si}}{p_s} $$
    is the generalized Kullback-Leibler divergence of order $q$ between each community and the meta-community.

    The generalized \(\beta\) entropy \({^\beta H_q}\) can be interpreted as the information gain produced by the knowledge of each local community's species probabilities related to the meta-community's probabilities. It is always positive and is independent of \(^\alpha H_q\) in the special case of \(q\rightarrow 1\), when the Shannon \(\beta\) entropy is obtained.

 

  • The \(\beta\) diversity of order \(q\) is then given by
    \( {^\beta D_q} = e_q\left({^\beta H_q}\right) \)
    \(\beta\) diversity is the equivalent number of communities, i.e. the number of equally-weighted, non-overlapping communities that would have the same diversity as the observed ones.

Partizionamento della Biodiversitá

  • Considerando una meta comunitá divisa in \(I\) comunitá locali indicate da \(i=1,\ldots,I\). Sia \(n_{si}\) il numero di esemplari campionati della specie \(s\) nella comunitá locale \(i\), con \(\sum_{s=1}^S n_{si}=n_i\) e \(\sum_{i=1}^I n_i=n\), dove \(n_i\) e \(n\) sono rispettivamente il numero di esemplari nella comunitá \(i\)-esima e nella meta-comunitá. In ogni comunitÁ \(i\), la probabilitá di un esemplare di apartenere alla specie \(s\) é calcolata con \(p_{si} =n_{si}/n_i\). La stessa probabilitá per la meta-comunitá é \(p_{s}\). Comunitá possono avere un peso \(w_i\), che soddisfa \(p_s = \sum_{i=1}^I w_ip_{si}\). \(w_{i}=n_{i}\) é comunemente usato come un possibile peso, ma la The commonly-used \(w_{i}=n_{i}\) is a possible weight, la scelta dei pesi puó essere arbitraria (es. le aree campionate).
  • L'entropia \(\gamma\) di Tsallis di ordine \(q\) della meta-comunitá é definata come :
    $$ ^\gamma H_q = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_s^q \right)=\cdots=-\sum_{s=1}^S p_s^q \ln_q p_s. $$
    L'entropia \(\alpha\) di Tsallis della comunitá \(i\)-esima é:
    $$ ^\alpha _i H_q = {1 \over q - 1} \left( 1 - \sum_{s=1}^S p_{si}^q \right). $$
    La definizione naturale dell'entropia \(\alpha\) totale della meta-comunitá é la media ponderata delle entropie \(\alpha\) delle comunitá locali
    $$ ^\alpha H_q = \sum_{i=1}^I w_i \;{^\alpha _i}H_q $$
  • I valori delle diversitá \(\alpha\) e \(\gamma\) sono dati dai numeri di Hill \(D_{q}\), ossia il numero delle specie "equi-frequenti" che daranno lo stesso livello di diversitá come i dati. Secondo la definizione di Routledge [3] [4] i numeri di Hill avranno la seguente espressione per le diversitá \(\gamma\) e \(\alpha\) della meta-comunitá:
    $$ ^\gamma D_q = \left(\sum_{s=1}^S p_s^q\right)^{\frac 1 {1-q}}=\cdots=e_q\left({^\gamma H_q}\right) $$

    $$ ^\alpha D_q = \left(\sum_{i=1}^I w_i \sum_{s=1}^S p_{si}^q\right)^{\frac 1 {q-1}} $$
  • I valori di diversitá sono tradizionalmente suddivise in diversitá \(\gamma\), \(\alpha\) e \(\beta\) [5], dove
    $$ {^\gamma D_q} ={^\alpha D_q} \times {^\beta D_q} $$
  • Partendo dalla partizione moltiplicativa della diversitá \(\gamma\), si ottiene la seguente media ponderata per l'entropia \(\beta\) generalizzata di ordine \(\q\):
    $$ {^\beta H_q} = {^\gamma H_q}-{^\alpha H_q} = \sum_{i=1}^I w_i\;{^\beta _i H_q} $$
    dove
    $$ {^\beta _i H_q}=\sum_{s=1}^S p_{si}^q \ln_q \frac{p_{si}}{p_s} $$
    é la divergenza generalizzata di kullback-Leibner di ordine \(\q\) fra le singole comunitá e la meta-comunitá.

    L'entropia generalizzata \(\beta\) \({^\beta H_q}\) puó essere interpretata come un guadagno di informazioni dovute alla conocscenza delle probabilitá delle specie delle comunitá locali in relazione alle probabilitá della meta-comunitá. É sempre positiva ed é indipendente da \(^\alpha H_q\) nel caso particolare di \(q\rightarrow 1\), quando si ottiene l'entropia \(\beta\) di Shannon.
  • La diversitá \(\beta\) di ordine \(\q\) é dataa quindi da
    \( {^\beta D_q} = e_q\left({^\beta H_q}\right) \)
    la diversitá \(\beta\) é il numero equivalente delle comunitá, ossia il numero delle comunitá "equi-ponderate" e non sovrapposte che avranno la stessa diversitá di quelle osservate.

Bias correction

  • Real data are almost always samples of larger communities, so some species may have been missed. The induced bias on Simpson entropy is smaller than on Shannon entropy because the former assigns lower weights to rare species, i.e. the sampling bias is even more important when \(q\) decreases.
  • The sample coverage of community \(i\), denoted \(C_i\), is the total probability of occurrence of the species observed in the sample. It is estimated from the number of singletons (species observed only once) of the sample, denoted \(S^1_i\), and the sample size \(n_i\):
    $$ C_i = 1 - \frac {S^1_i}{n_i} $$
    and for the meta-community
    $$ C = 1 - \frac {S^1}{n} $$
  • Unbiased estimates of the probabilities \(p_{si}\) and \(p_s\) are respectively given by \(\hat p_{si}=C_i n_{si}/n_i\) and \(\hat p_s=C n_{s}/n\)
  • Besides the unbiased estimates of the probabilities based on sample coverage, estimation bias corrections are also introduced for the definitions of entropy given in the previous section. Chao and Shen's correction relies on the Horvitz-Thomson estimator which corrects a sum of measurements for missing species by dividing each measurement by the probability for each species to be present in the \(i\)-th sample, i.e. by \(1-(1-p_{si})^{n_i}\).
  • Combining the two corrections, the following alternative expressions are obtained for \({^\gamma H_q}\) and \({^\beta_i H_q}\)
    $$ {^\gamma \hat H_q}=-\sum_{s=1}^S \frac{\hat p_s^q \ln_q \hat p_s}{1-(1-\hat p_{s})^n} $$
    $$ {^\beta_i \hat H_q}=\sum_{s=1}^S \frac{\hat p_{si}^q \ln_q \frac{\hat p_{si}}{\hat p_s}}{1-(1-\hat p_{si})^{n_i}} $$
  • Another estimation bias is due to the non-linearity of entropy measures. Notice that as probabilities \(p_s\) are estimated by \(n_s/n\), estimating \(p^q_s\) by \(n_s^q/n^q\) (for \(q>0\)) is an important source of underestimation of entropy. Grassberger [6], Holste et al. [8], Hou et al. [9] and Bonachela et al. [10] derived alternative unbiased estimators of \(p_s^q\) to be plugged into the formula of the Tsallis \(\gamma\) entropy.
  • The correction for missing species by Chao and Shen and that for non-linearity by Grassberger ignore each other. Chao and Shen's bias correction is important when \(q\) is small and becomes negligible for \(q=2\) while Grassberger's correction increases with \(q\), vanishing for \(q=0\). A rough but pragmatic estimation-bias correction is the maximum value of the two corrections.

Correzione del bias

  • i dati reali sono quasi sempre campioni di comunitá piú ampie, cosí alcune specie potrebbero essere andate perse. L'errore sull'entropia di Simpson é minore di quello sull'entropia di Shannon perché la prima assegna pesi minori a specie rare, ossia l'errore di campionamento é importante quando \(q\) diminuisce.
  • La copertura campionaria della comunitá \(i\), indicata con \(C_i\), é la probabilitá di presenze delle specie osservate nel campioni. É data dal numero di elementi singoli (specie osservate solo una volta) del campione, indicato con \(S^1_i\), e la dimensione del campione \(n_i\):
    $$ C_i = 1 - \frac {S^1_i}{n_i} $$
    e per la meta-comunitá
    $$ C = 1 - \frac {S^1}{n} $$
  • Le valutazioni eque delle probabilitá \(p_{si}\) e \(p_s\) sono date rispettivamente da \(\hat p_{si}=C_i n_{si}/n_i\) e \(\hat p_s=C n_{s}/n\)
  • Accanto alle "valutazioni eque delle probabilitá" basate su coperture campionare, le correzioni degli errori di stima sono anche introdotte per la definizione di entropia data nella sezione precedente. La correzione di Chao e Shen si basa sullo stimatore di Horvitz-Thomson che corregge una somma di misure per una specie assente dividendo ogni misura per la probabilitá che ogni specie ha di essere presente nel campione \(i\)-esimo, ossia diviso per \(1-(1-p_{si})^{n_i}\).
  • Combinando i due tipi di correzioni, si ottengono le seguenti espressioni alternative per \({^\gamma H_q}\) e \({^\beta_i H_q}\)
    $$ {^\gamma \hat H_q}=-\sum_{s=1}^S \frac{\hat p_s^q \ln_q \hat p_s}{1-(1-\hat p_{s})^n} $$
    $$ {^\beta_i \hat H_q}=\sum_{s=1}^S \frac{\hat p_{si}^q \ln_q \frac{\hat p_{si}}{\hat p_s}}{1-(1-\hat p_{si})^{n_i}} $$
  • Un'altro errore di stima é dovuto alla non-linearitá delle misure dell'entropiá. Si noti che siccome le probabilitá \(p_s\) sono stimate da \(n_s/n\), valutando \(p^q_s\) da \(n_s^q/n^q\) (per \(q>0\)) é una fonte importante per la sottovalutazione dell'entropia. Grassberger [6], Holste et al. [8], Hou et al. [9] and Bonachela et al. [10] hanno ricavato un valutatore alternativo di \(p_s^q\) da inserire nella formula dell'entropia \(\gamma\) di Tsallis.
  • La correzione delle specie mancanti di Chao e Shen e quella per la non-linearitá di Grassberger si ingorano a vicenda. La correzione di Chao e Shen é importante quando \(q\) é piccolo e diventa trascurabile per \(q=2\) mentre la correzione di Grassberger aumenta con \(q\), svanendo per \(q=0\). Una correzione di massima, ma pragmatica, dell'errore di stima é quella di prendere il massimo valore delle due correzioni.

References

[1] Tsallis C. (1988). Possible generalization of boltzmann-gibbs statistics. Journal of Statistical Physics, 52, 479-487.

[2] Marcon E., Scotti I., Hérault B., Rossi V., Lang G. (2014). Generalization of the partitioning of Shannon diversity. PLoS ONE 9(3): e90289.doi:10.1371/journal.pone.0090289.

[3] Routledge R. (1979). Diversity indices: Which ones are admissible? Journal of Theoretical Biology, 76, 503-515.

[4] Patil G. P., Taillie C. (1982). Diversity as a concept and its measurement.Journal of the American Statistical Association, 77, 548-561.

[5] Whittaker, R. H. (1960). Vegetation of the Siskiyou Mountains, Oregon and California. Ecological Monographs, 30, 279-338.

[6] Grassberger P. (1988). Finite sample corrections to entropy and dimension estimates. Physics Letters A, 128, 369-373.

[7] Hill M. O. (1973). Diversity and Evenness: A Unifying Notation and Its Consequences. Ecology 54, 427-432.

[8] Holste D., Grobe I., Herzel H. (1998). Bayes' estimators of generalized entropies. Journal of Physics A: Mathematical and General, 31, 2551-2566.

[9] Hou Y., Wang B., Song D., Cao X., Li W. (2014). Quadratic tsallis entropy bias and generalized maximum entropy models. Computational Intelligence, 30, 233-262.

[10] Bonachela J. A., Hinrichsen H., Muñoz M. A. (2008). Entropy estimates of small data sets. Journal of Physics A: Mathematical and Theoretical, 41,1-9.

Riferimenti

[1] Tsallis C. (1988). Possible generalization of boltzmann-gibbs statistics. Journal of Statistical Physics, 52, 479-487.

[2] Marcon E., Scotti I., Hérault B., Rossi V., Lang G. (2014). Generalization of the partitioning of Shannon diversity. PLoS ONE 9(3): e90289.doi:10.1371/journal.pone.0090289.

[3] Routledge R. (1979). Diversity indices: Which ones are admissible? Journal of Theoretical Biology, 76, 503-515.

[4] Patil G. P., Taillie C. (1982). Diversity as a concept and its measurement.Journal of the American Statistical Association, 77, 548-561.

[5] Whittaker, R. H. (1960). Vegetation of the Siskiyou Mountains, Oregon and California. Ecological Monographs, 30, 279-338.

[6] Grassberger P. (1988). Finite sample corrections to entropy and dimension estimates. Physics Letters A, 128, 369-373.

[7] Hill M. O. (1973). Diversity and Evenness: A Unifying Notation and Its Consequences. Ecology 54, 427-432.

[8] Holste D., Grobe I., Herzel H. (1998). Bayes' estimators of generalized entropies. Journal of Physics A: Mathematical and General, 31, 2551-2566.

[9] Hou Y., Wang B., Song D., Cao X., Li W. (2014). Quadratic tsallis entropy bias and generalized maximum entropy models. Computational Intelligence, 30, 233-262.

[10] Bonachela J. A., Hinrichsen H., Muñoz M. A. (2008). Entropy estimates of small data sets. Journal of Physics A: Mathematical and Theoretical, 41,1-9.