Skip to content

chore(bigtable): run samples tests pre-submit - #18599

Open
daniel-sanche wants to merge 12 commits into
mainfrom
samples_1_run_tests
Open

daniel-sanche wants to merge 12 commits into
mainfrom
samples_1_run_tests

Conversation

@daniel-sanche

@daniel-sanche daniel-sanche commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

The bigtable library lost the test configs for samples as part of the move to the monorepo, so they are left untested. This PR adds back the configs, and runs tests as part of the pre-submit check

Samples tests are now run from the central bigtable noxfile, instead of using individual noxfiles for each sample directory

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates relative imports across several Google Cloud Bigtable samples and adds new configuration files, including noxfile.py and requirements files, for the async data client samples. A critical issue was identified in packages/google-cloud-bigtable/samples/utils.py, where unresolved git merge conflict markers were left in the code, which will lead to a runtime SyntaxError.

Comment thread packages/google-cloud-bigtable/samples/utils.py Outdated
@daniel-sanche
daniel-sanche added this pull request to stack #18600 October 7, 2026 23:48
@daniel-sanche daniel-sanche changed the title [DRAFT] chore(bigtable): run samples tests pre-submit chore(bigtable): run samples tests pre-submit Oct 9, 2026
@daniel-sanche
daniel-sanche marked this pull request as ready for review October 9, 2026 01:04
@daniel-sanche
daniel-sanche requested review from a team as code owners October 9, 2026 01:04
@parthea
parthea force-pushed the samples_1_run_tests branch from 1cc22cd to 886ab68 Compare October 9, 2026 16:19
@parthea

parthea commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

parthea force-pushed

To clarify , I clicked the rebase stack button

@parthea parthea self-assigned this Oct 9, 2026
apache-beam==2.69.0; python_version == '3.9'
apache-beam==2.71.0; python_version >= '3.10'
google-cloud-bigtable==2.35.0
google-cloud-bigtable

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please could you add a comment to clarify the reason we don't pin the google-cloud-bigtable library here, and elsewhere?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We want these sample tests to run against the repo head, so we can catch if any of our changes break anything. Running against an old released version of the library isn't valuable

@parthea parthea Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, perhaps we can remove the requirements files altogether, or only have the dependencies which are not in setup.py

We can then remove session.install("-e", ".", "--no-deps") from the noxfile.py file and have:

    session.install("-e", ".")
    session.install(
        "google-cloud-testutils",
        "mock",
        "pytest",
        "pytest-asyncio",
        *req_args,
    )

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the requrements are useful, even if just to help users understand/run the samples. They're referenced in the samples README, and could be seen as of part of the documentation

But if you prefer to get rid of them, I don't feel too strongly about it

Comment on lines +69 to +76
batcher = table.mutations_batcher(flush_count=2)
rows = table.read_rows()
for row in rows:
row = table.row(row.row_key)
row = table.direct_row(row.row_key)
row.delete_cell(column_family_id="cell_plan", column="data_plan_01gb")

batcher.mutate_rows(rows)
batcher.close()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Gemini suggested we should have this instead

change

    batcher = table.mutations_batcher(flush_count=2)
    rows = table.read_rows()
    for row in rows:
        row = table.direct_row(row.row_key)
        row.delete_cell(column_family_id="cell_plan", column="data_plan_01gb")

    batcher.mutate_rows(rows)
    batcher.close()

to

    with table.mutations_batcher(flush_count=2) as batcher:
        for row in table.read_rows():
            direct_row = table.direct_row(row.row_key)
            direct_row.delete_cell(column_family_id="cell_plan", column="data_plan_01gb")
            batcher.mutate(direct_row)

This test fails without the fix

def test_streaming_and_batching_actually_deletes(table_id):
    """Verifies that streaming_and_batching actually deletes cells from the table."""
    from google.cloud import bigtable
    from . import deletes_snippets

    client = bigtable.Client(project=PROJECT, admin=True)
    instance = client.instance(BIGTABLE_INSTANCE)
    table = instance.table(table_id)

    # 1. Seed a row in the table with cell_plan:data_plan_01gb
    row_key = b"phone#4c410523#20190501"
    row = table.direct_row(row_key)
    row.set_cell("cell_plan", b"data_plan_01gb", b"true")
    row.commit()

    # 2. Verify row exists before calling the snippet
    seeded_row = table.read_row(row_key)
    assert seeded_row is not None, "Failed to seed row!"
    assert b"data_plan_01gb" in seeded_row.cells["cell_plan"]

    # 3. Run the snippet
    deletes_snippets.streaming_and_batching(
        PROJECT, BIGTABLE_INSTANCE, table_id
    )

    # 4. Check if the cell was deleted
    updated_row = table.read_row(row_key)
    # If the snippet worked, either the row is None (all cells deleted)
    # or the column data_plan_01gb is absent from cell_plan:
    has_cell = (
        updated_row is not None
        and b"data_plan_01gb" in updated_row.cells.get("cell_plan", {})
    )
    assert not has_cell, "Cell was NOT deleted because batcher received an exhausted generator!"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, fixed!

@parthea parthea assigned daniel-sanche and unassigned parthea Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants