Detailansicht
A comparison of the logistic regression model and the random coefficients linear mixed model for analysis of early kidney disease progression
Maria Neubauer
Art der Arbeit
Magisterarbeit
Universität
Universität Wien
Fakultät
Fakultät für Wirtschaftswissenschaften
Betreuer*in
Georg Heinze
URN
urn:nbn:at:at-ubw:1-29861.16621.421666-0
Link zu u:search
(Print-Exemplar eventuell in Bibliothek verfügbar)
Abstracts
Abstract
(Deutsch)
Die Früherkennung der chronischen Nierenerkrankung (engl. CKD), eines der größten aktuellen Gesundheitsprobleme, ist essentiell. Die Albuminausscheidungsrate durch den Urin (engl. UACR) ist ein kontinuierlicher klinischer Marker für Progression der CKD. Folglich gibt es zwei mögliche Ansätze um die Progressionswahrscheinlichkeit eines Patienten/einer Patientin zu schätzen: entweder durch eine logistische Regression, in der eine binäre abhängige Variable verwendet wird, welche den Status eines Patienten/einer Patientin (progredient oder stabil) angibt; oder durch eine longitudinale Herangehensweise, in welcher die wiederholten Messungen von UACR als abhängige Variable verwendet werden. Das Ziel dieser Magisterarbeit ist nun diese beiden Modelle, die logistische Regression und das longitudinale random coefficients model, in der Analyse der Progression der CKD zu vergleichen. Die Modelle werden dabei in Bezug auf ihre statistische Mächtigkeit, den Effekt eines Risikofaktors (oder einer Behandlung) zu sichern, und auf ihre Genauigkeit in der Schätzung von Progressionsraten verglichen. Es wird gezeigt, wie statistische Effektstärkenschätzer, die üblicherweise nur für binäre abhängige Variablen zur Verfügung stehen, wie z. B. erwartetes relatives Risiko, erwartete Odds Ratio, concordance statistics und integrated discrimination improvements, in beiden Ansätzen anhand simulierter Datensätze und eines Praxisbeispiels geschätzt werden können. Alle Berechnungen werden in zwei verschiedenen Situationen durchgeführt: unter Verwendung von log UACR (LOG Modelle) und nach einer Box-Cox Transformation von log UACR (BCLOG Modelle).
Die Simulation und die Analyse des ONTARGET-SysKid Datensatzes haben beide klare Unterschiede zwischen den LOG Modellen und den BCLOG Modellen gezeigt. Da die Annahme normalverteilter Residuen in den longitudinalen LOG Modellen verletzt war, unterschätzten diese Modelle die Progressionswahrscheinlichkeit. Aufgrund ihrer höheren Genauigkeit und Unverzerrtheit sollten die longitudinalen BCLOG Modelle die erste Wahl in die Analyse sein.
Die longitudinalen BCLOG Modelle erreichten bessere Ergebnisse als die logistischen BCLOG Modelle: In der Simulation war die geschätzte Progressionswahrscheinlichkeit in den longitudinalen Modellen immer näher an der wahren Progressionswahrscheinlichkeit als die der logistischen Modelle, welche zudem eine höhere Variabilität aufwiesen. Die Power des Risikofaktors für Progression, welche stark von Stichprobenumfang abhängt, war in den longitudinalen Modellen immer höher als in den logistischen. Die Analyse der Modellverbesserung zeigte nur kleine Unterschiede zwischen den logistischen und longitudinalen Modellen.
Wir fanden einen großen Unterschied zwischen den tatsächlichen Progressionsraten und der wahren Progressionswahrscheinlichkeit, was daran liegt, dass PatientInnen mit einer UACR Messung über 33.9 mg/mmol bei Baseline von der Studie ausgeschlossen wurden. Wenn bei UACR ein Messfehler passiert, dann ist die Kategorisierung des kontinuierlichen Markers in Progression/nicht Progression verzerrt in Richtung höherer Progressionsraten. Das longitudinale Modell schätzt diesen Messfehler explizit und garantiert somit fast unverzerrte Progressionsraten.
Abstract
(Englisch)
Chronic kidney disease (CKD) is a major public health problem and therefore its early detection is crucial. Urine albumin to creatinine ratio (UACR) is a continuous clinical marker for the progression (= incidence or progression) of CKD. Consequently, there are two possible approaches to estimate the probability of progression for a patient: either by a logistic approach, using a binary outcome variable defining the status of a patient (progressing or stable), or by a longitudinal approach, using the repeated measurements of UACR as outcome variable. The aim of this Master thesis is to compare these two approaches for the analysis of early kidney disease progression. Therefore the models were compared with regard to their power to detect an association of a risk factor (or treatment) with progression, and with regard to their accuracy in estimating progression rates. It is shown how quantities usually available only with binary outcomes, such as expected relative risks, expected odds ratios, concordance statistics and integrated discrimination improvements can be estimated by either approach, in simulated datasets as well as in a real-life example. All calculations were done using the original log UACR (LOG models) or a Box-Cox log transformation thereof (BCLOG models).
The simulation and the real-life data analysis both showed small but clear differences of the misspecified LOG models compared to BCLOG models. Since the assumption of normally distributed residuals in the longitudinal LOG models was violated, the LOG models underestimated the probability of progression. Therefore, the longitudinal model on Box-Cox transformed log UACR should be the first choice for analysis because of its higher accuracy and unbiasedness.
The longitudinal BCLOG model attained better results than the logistic BCLOG model: the estimated probability of progression by the longitudinal model in the simulation was always closer to the true probability of progression than the probability of progression estimated by the logistic model, which moreover showed a higher variability. The power to detect an effect of a risk factor x on progression was always higher in the longitudinal model than in the logistic model. Testing different scenarios showed that the choice of sample size is crucial for achieving adequate power. Assessing model improvement only lead to small differences in the logistic and longitudinal models.
Furthermore, we noticed a marked difference between apparent progression rates and the true probability of progression, which was due to the fact that patients with an UACR above a certain cut-off point at baseline were excluded from this observational study. Therefore, if UACR is measured with error, then the assessment of progression by categorization of this underlying continuous marker is biased towards higher progression rates. The longitudinal model explicitly estimates this measurement error and therefore leads to approximately unbiased progression rates.
Schlagwörter
Schlagwörter
(Englisch)
logistic regression random coefficients linear mixed model longitudinal model dichotomization Box-Cox transformation /chronic kidney disease progression simulation ONTARGET-SysKid study
Schlagwörter
(Deutsch)
logistische Regression linear gemischtes Modell mit Zufallskoeffizienten longitudinales Modell Dichotomisierung Box-Cox Transformation chronische Nierenerkrankung Progression Simulation ONTARGET-SysKid Studie
Autor*innen
Maria Neubauer
Haupttitel (Englisch)
A comparison of the logistic regression model and the random coefficients linear mixed model for analysis of early kidney disease progression
Paralleltitel (Deutsch)
Vergleich zwischen dem logistischen Regressionsmodell und dem linear gemischten Modell mit Zufallskoeffizienten für die Analyse der Nierenerkrankungsprogression
Publikationsjahr
2014
Umfangsangabe
III, 115 S. : graph. Darst.
Sprache
Englisch
Beurteiler*in
Georg Heinze
Klassifikation
44 Medizin > 44.32 Medizinische Mathematik, medizinische Statistik
AC Nummer
AC12133693
Utheses ID
30826
Studienkennzahl
UA | 066 | 951 | |
