Combine rows with same id and delete duplicated rows

Question

After merging some data, I have multiple rows per ID. I ONLY want to keep multiple SAME ID's if the data differs. An NA value should be considered equal to any colwise data point.

data:

df <- structure(list(id = c(1L, 2L, 2L, 2L, 3L, 3L, 4L, 4L, 4L, 5L), 
    v1 = structure(c(1L, 1L, NA, 1L, 1L, 1L, 1L, NA, 1L, 1L), .Label = "a", class = "factor"), 
    v2 = structure(c(1L, 2L, 2L, 3L, 1L, 1L, 1L, 1L, NA, 1L), .Label = c("a", 
    "b", "c"), class = "factor"), v3 = structure(c(1L, 1L, 1L, 
    1L, 1L, 1L, NA, 2L, 2L, 1L), .Label = c("a", "b"), class = "factor")), .Names = c("id", 
"v1", "v2", "v3"), row.names = c(NA, -10L), class = "data.frame")

looks like:

   id   v1   v2   v3
    1    a    a    a
    2    a    b    a
    2     b    a
    2    a    c    a
    3    a    a    a
    3    a    a    a
    4    a    a 
    4     a    b
    4    a     b
    5    a    a    a

desired output:

   id   v1   v2   v3
    1    a    a    a
    2    a    b    a
    2    a    c    a
    3    a    a    a
    4    a    a    b
    5    a    a    a

Happy if there exists a data.table solution.

Jaap · Accepted Answer

A possible solution using the data.table-package:

library(data.table)
setDT(df)[, lapply(.SD, function(x) unique(na.omit(x))), by = id]

which gives:

   id v1 v2 v3
1:  1  a  a  a
2:  2  a  b  a
3:  2  a  c  a
4:  3  a  a  a
5:  4  a  a  b
6:  5  a  a  a

Combine rows with same id and delete duplicated rows

data:

looks like:

desired output:

Answers (2)

Related Questions