identify duplicate in data frame and mark them

--Hi,

i have this structure:

structure(list(sample = c("CTRLM", "CTRLM", "1", "1", "1", "1", 
"1", "SC1643", "SC1643", "H9 P18"), genes_ref1 = c("RPP30", "H2BC8", 
"SLAIN2", "SLAIN2", "SLAIN2", "SLAIN2", "SLAIN2", "RPP30", "H2BC8", 
"NECTIN2"), MPX = c("MPX4", "MPX4", "MPX1", "MPX2", "MPX2", "MPX3", 
"MPX4", "MPX4", "MPX4", "MPX4"), genes_ref2 = c("H2BC8", "RPP30", 
"RPPH1", "RPP30", "RPPH1", "RPP30", "RPPH1", "H2BC8", "RPP30", 
"RPPH1")), row.names = c(NA, -10L), class = "data.frame")

and i want to find duplicates on rows and obtain that:

  sample ref1  MPX ref2	doublon
1   CTRLM      RPP30 MPX4      H2BC8	true
2   CTRLM      H2BC8 MPX4      RPP30	true
3       1     SLAIN2 MPX1      RPPH1	false
4       1     SLAIN2 MPX2      RPP30	false
5       1     SLAIN2 MPX2      RPPH1	false
6       1     SLAIN2 MPX3      RPP30	false
7       1     SLAIN2 MPX4      RPPH1	false
8  SC1643      RPP30 MPX4      H2BC8	true
9  SC1643      H2BC8 MPX4      RPP30	true
10 H9 P18    NECTIN2 MPX4      RPPH1	false

if ref1==ref2 for the same MPX then it's a duplicate

Your criterion for what is a duplicate is a bit unclear. Suppose that ref2 in the first record were changed from H2BC8 to something unmatched in the data (say ABCD5). Would record 1 still be a duplicate (since its ref1 matches ref2 in record 2 with the same MPX)?

--Hi,

i found this solution:

make_key <- function(d, a, b) paste(d$sample, d$MPX, d[[a]], d[[b]], sep = "|")

key     <- make_key(df_temp, "genes_ref1", "genes_ref2")
rev_key <- make_key(df_temp, "genes_ref2", "genes_ref1")
df_temp$doublon <- rev_key %in% key & key != rev_key