GPandas

Grouping & Aggregation

Group rows by column values and compute aggregations, including multiple functions per column with Agg

Learn how to group a DataFrame by one or more columns and compute aggregations in GPandas. Use the built-in reducers (Mean, Sum, Min, Max) for quick summaries, Agg for multiple functions per column, or Apply for arbitrary per-group logic.

Overview

A group-by operation splits the DataFrame into groups, applies a function to each group, and combines the results:

OperationMethodDescription
GroupGroupBy()Split rows by one or more columns
ReduceMean(), Sum(), Min(), Max()Single aggregation across numeric columns
Multi-aggregateAgg()Multiple functions per column
CustomApply()Arbitrary function per group

GroupBy

Creates a grouped view of the DataFrame.

Function Signature

func (df *DataFrame) GroupBy(by []string, axis int) (*GroupBy, error)
ParameterDescription
byColumn names to group by
axisMust be 0 (group rows); other axes are not yet supported

Sample Data

All examples use this DataFrame:

DepartmentSalaryUnits
Eng1003
Sales505
Eng2002
Sales708
Eng1504

Setup Code

package main

import (
    "fmt"
    "log"

    "github.com/apoplexi24/gpandas"
    "github.com/apoplexi24/gpandas/dataframe"
)

func main() {
    gp := gpandas.GoPandas{}

    df, _ := gp.DataFrame(
        []string{"Department", "Salary", "Units"},
        []gpandas.Column{
            {"Eng", "Sales", "Eng", "Sales", "Eng"},
            {100.0, 50.0, 200.0, 70.0, 150.0},
            {int64(3), int64(5), int64(2), int64(8), int64(4)},
        },
        map[string]any{
            "Department": gpandas.StringCol{},
            "Salary":     gpandas.FloatCol{},
            "Units":      gpandas.IntCol{},
        },
    )

    // Examples follow...
}

Single Aggregations

The reducers compute one statistic across all numeric columns of each group, returning a new DataFrame:

gb, err := df.GroupBy([]string{"Department"}, 0)
if err != nil {
    log.Fatalf("GroupBy failed: %v", err)
}

means, _ := gb.Mean()  // also: Sum(), Min(), Max()
fmt.Println(means.String())

Each reducer returns a DataFrame with one row per group, ordered by group key.


Agg — Multiple Functions per Column

Agg applies one or more aggregation functions to specific columns, producing a column named <column>_<func> for each pair.

Function Signature

func (gb *GroupBy) Agg(spec map[string][]AggFunc) (*DataFrame, error)

Supported Functions

ConstantOperation
AggSumSum
AggMeanMean
AggCountCount of non-null values
AggMin / AggMaxMinimum / maximum
AggStdSample standard deviation (ddof=1)
AggMedianMedian
AggFirst / AggLastFirst / last non-null value

Note: Numeric functions ignore null and non-numeric values. AggCount counts non-null values; AggFirst/AggLast work with any type.

Example

gb, _ := df.GroupBy([]string{"Department"}, 0)
result, err := gb.Agg(map[string][]dataframe.AggFunc{
    "Salary": {dataframe.AggSum, dataframe.AggMean, dataframe.AggMax},
    "Units":  {dataframe.AggSum},
})
if err != nil {
    log.Fatalf("Agg failed: %v", err)
}
fmt.Println(result.String())

Output

+------------+------------+-------------+------------+-----------+
| Department | Salary_sum | Salary_mean | Salary_max | Units_sum |
+------------+------------+-------------+------------+-----------+
| Eng        | 450        | 150         | 200        | 9         |
| Sales      | 120        | 60          | 70         | 13        |
+------------+------------+-------------+------------+-----------+
[2 rows x 5 columns]

Aggregation Flow


Apply — Custom Per-Group Logic

Apply runs an arbitrary function on each group's sub-DataFrame and concatenates the results:

func (gb *GroupBy) Apply(f func(*DataFrame) (*DataFrame, error)) (*DataFrame, error)
gb, _ := df.GroupBy([]string{"Department"}, 0)
result, _ := gb.Apply(func(group *dataframe.DataFrame) (*dataframe.DataFrame, error) {
    // e.g. return the top earner per department
    return group.SortValues(dataframe.SortOptions{
        By:        []string{"Salary"},
        Ascending: []bool{false},
    })
})

Error Handling

Common Errors

ErrorCauseSolution
"column ... not found"Invalid group or value columnVerify the column exists
"axis ... is not supported"axis other than 0Use 0 to group rows
"spec must contain at least one column"Empty Agg specProvide at least one column

Thread Safety

Grouping reads the DataFrame under a read lock and produces new DataFrames, so the original is never mutated and concurrent grouping is safe.


See Also

On this page