GPandas

Unique Values & Deduplication

Find distinct values, count cardinality, and remove duplicate rows with Unique, NUnique, Duplicated, and DropDuplicates

Learn how to analyze cardinality and remove duplicate rows in GPandas. Unique and NUnique summarize distinct values, while Duplicated and DropDuplicates identify and remove repeated rows.

Overview

GPandas provides four methods for uniqueness and deduplication:

OperationMethodReturns
Distinct valuesUnique()[]any
Distinct countNUnique()int
Duplicate maskDuplicated()[]bool
Remove duplicatesDropDuplicates()*DataFrame

Sample Data

All examples use this DataFrame with repeated rows:

AB
x1
y2
x1
x9
z3

Setup Code

package main

import (
    "fmt"
    "log"

    "github.com/apoplexi24/gpandas/dataframe"
    "github.com/apoplexi24/gpandas/utils/collection"
)

func main() {
    a, _ := collection.NewStringSeriesFromData(
        []string{"x", "y", "x", "x", "z"}, nil)
    b, _ := collection.NewInt64SeriesFromData(
        []int64{1, 2, 1, 9, 3}, nil)

    df := &dataframe.DataFrame{
        Columns:     map[string]collection.Series{"A": a, "B": b},
        ColumnOrder: []string{"A", "B"},
        Index:       []string{"0", "1", "2", "3", "4"},
    }

    // Examples follow...
}

Unique and NUnique

Unique returns the distinct values of a column in order of first appearance. If the column contains nulls, a single nil entry is included at its first occurrence. NUnique returns the count of distinct non-null values.

Function Signatures

func (df *DataFrame) Unique(column string) ([]any, error)
func (df *DataFrame) NUnique(column string) (int, error)

Example

values, _ := df.Unique("A")
count, _ := df.NUnique("A")
fmt.Printf("Unique(A) = %v\n", values)
fmt.Printf("NUnique(A) = %d\n", count)

Output

Unique(A) = [x y z]
NUnique(A) = 3

Note: NUnique excludes nulls, matching pandas' default. Unique includes a single nil entry when the column has nulls.


Duplicated

Returns a boolean slice marking duplicate rows, aligned to row order. Rows are compared on the columns in subset (or all columns if subset is empty).

Function Signature

func (df *DataFrame) Duplicated(subset []string, keep string) ([]bool, error)

Keep Modes

keepMarked as duplicate (true)
"first" (default)All occurrences except the first
"last"All occurrences except the last
"none"All occurrences of any duplicated row

Example

Comparing on column A (values x, y, x, x, z):

first, _ := df.Duplicated([]string{"A"}, "first")
last, _  := df.Duplicated([]string{"A"}, "last")
none, _  := df.Duplicated([]string{"A"}, "none")
fmt.Printf("first: %v\n", first)
fmt.Printf("last:  %v\n", last)
fmt.Printf("none:  %v\n", none)

Output

first: [false false true true false]
last:  [true false true false false]
none:  [true false true true false]

The three x rows are at indices 0, 2, 3. With "first", indices 2 and 3 are flagged; with "last", indices 0 and 2; with "none", all three.


DropDuplicates

Returns a new DataFrame with duplicate rows removed, comparing on subset (or all columns). Index labels of the surviving rows are preserved.

Function Signature

func (df *DataFrame) DropDuplicates(subset []string, keep string) (*DataFrame, error)

Keep Modes

keepBehaviour
"first" (default)Keep the first occurrence
"last"Keep the last occurrence
"none"Drop all rows that have duplicates

Single-Column Subset

deduped, _ := df.DropDuplicates([]string{"A"}, "first")
fmt.Println(deduped.String())
+---+---+
| A | B |
+---+---+
| x | 1 |
| y | 2 |
| z | 3 |
+---+---+
[3 rows x 2 columns]

Multi-Column Subset

Considering both A and B, only (x, 1) is repeated, so just one row is removed:

deduped, _ := df.DropDuplicates([]string{"A", "B"}, "first")
fmt.Println(deduped.String())
+---+---+
| A | B |
+---+---+
| x | 1 |
| y | 2 |
| x | 9 |
| z | 3 |
+---+---+
[4 rows x 2 columns]

Deduplication Flow

Note: Row keys handle nulls explicitly, so null values never collide with real values during comparison.


Error Handling

Common Errors

ErrorCauseSolution
"DataFrame is nil"Operating on nil DataFrameCheck DataFrame initialization
"column 'X' not found"Invalid column or subset entryVerify the column exists
"keep must be 'first', 'last', or 'none'"Invalid keep valueUse a valid keep mode

Thread Safety

Uniqueness operations are thread-safe and read-only:

MethodLock TypeDescription
Unique() / NUnique()RLockRead lock during scan
Duplicated()RLockRead lock during key building
DropDuplicates()RLockRead lock during evaluation

Complete Example: Cleaning a Contact List

package main

import (
    "fmt"
    "log"

    "github.com/apoplexi24/gpandas"
)

func main() {
    gp := gpandas.GoPandas{}

    df, err := gp.Read_csv("contacts.csv")
    if err != nil {
        log.Fatalf("Failed to load data: %v", err)
    }

    // How many distinct email domains?
    count, err := df.NUnique("Domain")
    if err != nil {
        log.Fatalf("NUnique failed: %v", err)
    }
    fmt.Printf("Distinct domains: %d\n", count)

    // Remove duplicate contacts by email, keeping the first occurrence
    deduped, err := df.DropDuplicates([]string{"Email"}, "first")
    if err != nil {
        log.Fatalf("DropDuplicates failed: %v", err)
    }

    fmt.Printf("Rows before: %d, after dedup: %d\n", df.Len(), deduped.Len())
}

See Also

On this page