select unique values with equal probability

Question

I have a data frame like the following

I want to get unique c1 values, where c2 can be chosen with equal probability if there are multiple rows with the same c1 value. For example, the final result can be:

c1 c2
1 2
2 2
3 2
...

"A random choice of c2 for each possible value of c1" is what I want.

Stefan Wager · Accepted Answer

Here's a simple way to do it. Let's say your dataframe is called df.

x = unique(df$c1);
y = sapply(x, function(arg)sample(df$c2[df$c1 == arg], 1));
new_df = data.frame(c1 = x, c2 = y);

select unique values with equal probability

Answers (2)

Related Questions