Skip to content

Subject: Question about uniqueness_gtest / gtest.standard with multi-group designs (ftmsRanalysis 1.1.0) #38

Description

@ycsong-pnnl

I have recently used ftmsRanalysis (v1.1.0) for G-tests of uniqueness on FT-ICR MS data,
and I would like to understand two behaviours in gtest.standard(). I would appreciate your
view on whether they're intended, and the reasoning if so.

Workflow

  peakObj <- as.peakData(edata, fdata, emeta, ..., data_scale = "pres")
  peakObj <- group_designation(peakObj, main_effects = c("Layer", "Respiration"))
      # -> 4 groups: BTM_High, BTM_Low, TOP_High, TOP_Low
  comp <- divideByGroupComparisons(peakObj,
                                   comparisons = list(c("BTM_High", "BTM_Low")))[[1]]
  res  <- summarizeGroupComparisons(comp,
            summary_functions = "uniqueness_gtest",
            summary_function_params = list(uniqueness_gtest = list(
              pres_fn = "nsamps", pres_thresh = 2, pvalue_thresh = 0.05)))

What I observe

After divideByGroupComparisons() extracts a two-group comparison, Group is still a
factor with all four levels. In gtest.standard():

grps = as.factor(group)
num.grps = length(levels(grps)) # 4
...
df = num.grps - 1 # 3
major.group[j] = ifelse(obs.miss[2, 1] > obs.miss[2, 2],
grp.names[1], grp.names[2])

  1. Degrees of freedom: a two-group comparison is tested with df = 3. If only the
    two levels being compared are kept, df = 1. The G statistic is the same either
    way; only the p-value changes.

  2. Major group: major.group comes from factor levels 1 and 2 (here BTM_High and
    BTM_Low) whatever pair is compared. For a pair that doesn't include those levels
    (e.g. TOP_High vs TOP_Low), the result names a group outside the comparison, and
    uniqueness_gtest records it as NA. In our data, TOP_High vs TOP_Low gives 0
    compounds unique to either group, while 2,350 compounds have p <= 0.05.

Minimal example

One compound present in all 11 samples of group A and absent from all 10 of group B:

  library(ftmsRanalysis)
  x   <- matrix(c(rep(1, 11), rep(NA, 10)), nrow = 1)
  lv  <- c("BTM_High", "BTM_Low", "TOP_High", "TOP_Low")
  g_a <- factor(c(rep("BTM_High", 11), rep("BTM_Low", 10)), levels = lv)
  g_b <- factor(c(rep("TOP_High", 11), rep("TOP_Low", 10)), levels = lv)

  ftmsRanalysis:::gtest.standard(x, g_a)
  #   pvals = 2.17e-06, major.group = "BTM_High"
  ftmsRanalysis:::gtest.standard(x, droplevels(g_a))
  #   pvals = 7.0e-08,  major.group = "BTM_High"
  ftmsRanalysis:::gtest.standard(x, g_b)
  #   pvals = 2.17e-06, major.group = "BTM_Low"

Questions

  1. Is the df meant to reflect the number of groups in the full design rather than
    in the comparison? If so, could you explain the reasoning?
  2. Is major.group meant to come from factor levels 1 and 2 regardless of which pair
    is compared?
  3. If the two-group behaviour is intended, is applying droplevels() to Group before
    the test the recommended approach? And has this changed in any release after
    1.1.0?

Thanks very much for your time, and for the package.

Best regards,
Young

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions