Your question is Debugging Data Code. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You're a Statistician at MSD, reviewing a colleague's helper function before it goes into the trial-reporting pipeline. It's meant to summarize a biomarker by dose group, including the placebo (0 mg) arm as its own group for comparison.
import pandas as pd
def summarize_trial(df, dose_col="dose_mg", result_col="biomarker_level", outliers=[]):
outliers.append(None)
df = df[df[dose_col] != 0]
groups = df.groupby(dose_col)
summary = pd.DataFrame()
for name, group in groups:
row = {
"dose_mg": name,
"n": len(group),
"mean": group[result_col].mean(),
"std": group[result_col].std(),
}
summary = summary.append(row, ignore_index=True)
summary["is_significant"] = summary["mean"] == summary["mean"].max()
return summary
Walk through how you'd debug this, what breaks and when, and what you'd change before it's trusted for a trial report.