Biodiversity partitioning - Partizionamento della biodiversità

Bias correction

  • Real data are almost always samples of larger communities, so some species may have been missed. The induced bias on Simpson entropy is smaller than on Shannon entropy because the former assigns lower weights to rare species, i.e. the sampling bias is even more important when \(q\) decreases.
  • The sample coverage of community \(i\), denoted \(C_i\), is the total probability of occurrence of the species observed in the sample. It is estimated from the number of singletons (species observed only once) of the sample, denoted \(S^1_i\), and the sample size \(n_i\):
    $$ C_i = 1 - \frac {S^1_i}{n_i} $$
    and for the meta-community
    $$ C = 1 - \frac {S^1}{n} $$
  • Unbiased estimates of the probabilities \(p_{si}\) and \(p_s\) are respectively given by \(\hat p_{si}=C_i n_{si}/n_i\) and \(\hat p_s=C n_{s}/n\)
  • Besides the unbiased estimates of the probabilities based on sample coverage, estimation bias corrections are also introduced for the definitions of entropy given in the previous section. Chao and Shen's correction relies on the Horvitz-Thomson estimator which corrects a sum of measurements for missing species by dividing each measurement by the probability for each species to be present in the \(i\)-th sample, i.e. by \(1-(1-p_{si})^{n_i}\).
  • Combining the two corrections, the following alternative expressions are obtained for \({^\gamma H_q}\) and \({^\beta_i H_q}\)
    $$ {^\gamma \hat H_q}=-\sum_{s=1}^S \frac{\hat p_s^q \ln_q \hat p_s}{1-(1-\hat p_{s})^n} $$
    $$ {^\beta_i \hat H_q}=\sum_{s=1}^S \frac{\hat p_{si}^q \ln_q \frac{\hat p_{si}}{\hat p_s}}{1-(1-\hat p_{si})^{n_i}} $$
  • Another estimation bias is due to the non-linearity of entropy measures. Notice that as probabilities \(p_s\) are estimated by \(n_s/n\), estimating \(p^q_s\) by \(n_s^q/n^q\) (for \(q>0\)) is an important source of underestimation of entropy. Grassberger [6], Holste et al. [8], Hou et al. [9] and Bonachela et al. [10] derived alternative unbiased estimators of \(p_s^q\) to be plugged into the formula of the Tsallis \(\gamma\) entropy.
  • The correction for missing species by Chao and Shen and that for non-linearity by Grassberger ignore each other. Chao and Shen's bias correction is important when \(q\) is small and becomes negligible for \(q=2\) while Grassberger's correction increases with \(q\), vanishing for \(q=0\). A rough but pragmatic estimation-bias correction is the maximum value of the two corrections.

Correzione del bias

  • i dati reali sono quasi sempre campioni di comunitá piú ampie, cosí alcune specie potrebbero essere andate perse. L'errore sull'entropia di Simpson é minore di quello sull'entropia di Shannon perché la prima assegna pesi minori a specie rare, ossia l'errore di campionamento é importante quando \(q\) diminuisce.
  • La copertura campionaria della comunitá \(i\), indicata con \(C_i\), é la probabilitá di presenze delle specie osservate nel campioni. É data dal numero di elementi singoli (specie osservate solo una volta) del campione, indicato con \(S^1_i\), e la dimensione del campione \(n_i\):
    $$ C_i = 1 - \frac {S^1_i}{n_i} $$
    e per la meta-comunitá
    $$ C = 1 - \frac {S^1}{n} $$
  • Le valutazioni eque delle probabilitá \(p_{si}\) e \(p_s\) sono date rispettivamente da \(\hat p_{si}=C_i n_{si}/n_i\) e \(\hat p_s=C n_{s}/n\)
  • Accanto alle "valutazioni eque delle probabilitá" basate su coperture campionare, le correzioni degli errori di stima sono anche introdotte per la definizione di entropia data nella sezione precedente. La correzione di Chao e Shen si basa sullo stimatore di Horvitz-Thomson che corregge una somma di misure per una specie assente dividendo ogni misura per la probabilitá che ogni specie ha di essere presente nel campione \(i\)-esimo, ossia diviso per \(1-(1-p_{si})^{n_i}\).
  • Combinando i due tipi di correzioni, si ottengono le seguenti espressioni alternative per \({^\gamma H_q}\) e \({^\beta_i H_q}\)
    $$ {^\gamma \hat H_q}=-\sum_{s=1}^S \frac{\hat p_s^q \ln_q \hat p_s}{1-(1-\hat p_{s})^n} $$
    $$ {^\beta_i \hat H_q}=\sum_{s=1}^S \frac{\hat p_{si}^q \ln_q \frac{\hat p_{si}}{\hat p_s}}{1-(1-\hat p_{si})^{n_i}} $$
  • Un'altro errore di stima é dovuto alla non-linearitá delle misure dell'entropiá. Si noti che siccome le probabilitá \(p_s\) sono stimate da \(n_s/n\), valutando \(p^q_s\) da \(n_s^q/n^q\) (per \(q>0\)) é una fonte importante per la sottovalutazione dell'entropia. Grassberger [6], Holste et al. [8], Hou et al. [9] and Bonachela et al. [10] hanno ricavato un valutatore alternativo di \(p_s^q\) da inserire nella formula dell'entropia \(\gamma\) di Tsallis.
  • La correzione delle specie mancanti di Chao e Shen e quella per la non-linearitá di Grassberger si ingorano a vicenda. La correzione di Chao e Shen é importante quando \(q\) é piccolo e diventa trascurabile per \(q=2\) mentre la correzione di Grassberger aumenta con \(q\), svanendo per \(q=0\). Una correzione di massima, ma pragmatica, dell'errore di stima é quella di prendere il massimo valore delle due correzioni.