【问题标题】:xlwings find specific char text start and end position and format itxlwings 查找特定的字符文本开始和结束位置并格式化
【发布时间】:2022-09-27 14:17:17
【问题描述】:

我有多个excel文件,每个文件有12张。

所以,在每张纸上,我都有一个固定的文本,如下所示 - “项目已被阻止”

所以,我想做以下

a)在出现的任何地方找到文本“项目已被阻止”并将其更改为如下格式(带有粗体红色),如下所示

b) 将 excel 文件另存为 .xlsx

我尝试了以下

req_text = \"Project has been blocked\"

for a_cell in ws.used_range:
        if a_cell.value == req_text:
            print(a_cell.address)
            col = a_cell.address[0]
            ws[col].characters.font.bold = True  #how to get the start and end position of my text
            ws[col].characters.font.color = (255, 0, 0)

但这不能正常工作。因为我无法获得文本的开始和结束位置。

我希望我的输出如下

  • 文本“项目已被阻止”的 6 个实例都在一个单元格中,对吗?
  • 是的,正确的(在这个例子中)。在一个单元格中,我们有多个相同关键字的副本。但在实时,它们也可以以相同的方式为另一个用户(另一行)重复。因此,无论它出现在哪里,我们都应该更改格式
  • 但是,是的,每一行(用户)只会在一个单元格中包含多个文本实例。
  • @moken - 哦,是的。谢谢莫肯。我会尽力让你知道。

标签: python excel dataframe formatting xlwings


【解决方案1】:

我已更改代码以在工作表的已用数据中包含对文本的 Excel 搜索,然后根据原始代码的需要使用 Bold-Red 更新该单元格文本。
我最终不得不为 Excel 搜索使用 while 循环,并在搜索循环回到第一个找到的单元格时中断。因此,代码会跟踪first_search_cell与 while 循环中找到的下一个单元格进行比较。
我将 Excel 搜索变量保留为常量,因此如果您想更改搜索选项,您就知道名称和值是什么。显然你可以删除你不想要的或者使用 import from Xlwings 常量。
否则它几乎是一样的。

...
# Excel Search constants
# class LookAt:
xlPart = 2  # from enum XlLookAt
xlWhole = 1  # from enum XlLookAt
# class FindLookIn:
xlComments = -4144  # from enum XlFindLookIn
xlFormulas = -4123  # from enum XlFindLookIn
xlValues = -4163  # from enum XlFindLookIn
# class SearchOrder:
xlByColumns = 2  # from enum XlSearchOrder
xlByRows = 1  # from enum XlSearchOrder
# class SearchDirection:
xlNext = 1  # from enum XlSearchDirection
xlPrevious = 2  # from enum XlSearchDirection


def find_next_cell(start_cell):
    found_cell = ws.api.UsedRange.Find(req_text,
                                       After=start_cell,
                                       LookIn=xlValues,
                                       SearchOrder=xlByRows,
                                       SearchDirection=xlNext,
                                       MatchCase=False)
    return found_cell


wb = xw.Book('foo.xlsx')
ws = wb.sheets('Sheet1')

req_text = "Project has been blocked"

# First cell to start searching for req_text
search_from_cell = ws.api.Range('A1')

count = 0
first_search_cell = ''
while True:
    # Search for next cell to update
    update_cell = find_next_cell(search_from_cell)

    # Excel search will restart search again from the beginning after the last match
    # is found exit the loop when find the first match again
    if update_cell._inner.Address != first_search_cell:
        print(update_cell._inner.Address)
        # Set the address of the first found cell
        if count == 0:
            first_search_cell = update_cell._inner.Address

        cell_column = update_cell._inner.Column
        cell_row = update_cell._inner.Row

        text = ws.range(cell_row, cell_column).value
        len_req_text = len(req_text)

        # Create a List of the start position for all instances of the req_text
        # tsi = text position index
        tsi_list = [index for index in range(len(text)) if text.startswith(req_text, index)]

        # Iterate the tsi list
        for i in range(len(tsi_list)):
            # Get the index of the text position, tps = text position start
            tps = tsi_list[i]
            # Use the tps as start of the character position of the req_text
            # and (tps + length of req_text) for the end character position
            ws.range(cell_row, cell_column).characters[tps:tps + len_req_text].font.bold = True
            ws.range(cell_row, cell_column).characters[tps:tps + len_req_text].font.color = (255, 0, 0)

        search_from_cell = ws.api.Range(update_cell._inner.Address.replace('$', ''))

        count += 1

    else:
        break

wb.save('foo.xlsx')
...

【讨论】:

  • 我们如何获得X 和Y 的值?因为它们在每张纸上出现超过 40-50 次。那么,我们如何得到它们的 X 和 Y?
  • 意思是,没有我可以知道它们出现在哪里的模式/规则。它们可能出现在任何行和列中。
  • 我想我应该将text.startswith 更改为text.contains?因为如示例所示,它们可能出现在随机关键字/文本之后。那么,您认为contains 合适吗?或者通过使用startswith,我们只是在我们的数据中寻找特定的字符位置
  • 非常感谢您的帮助。如果我们没有contains 属性,您认为我们能够正确识别req_text 吗?因为它们总是只出现在某些关键字/文本之后(单元格文本不以“项目已被阻止”开头)。因此,单元格文本通常包含req_text,但不以req_text 开头
  • 以。。开始只是说这个位置的文本是否与“read_text”文本匹配。我们正在指定开始改变我们将文本开头的位置。所以这就像在'req_text'存在的点处将文本字符串切碎。我可能会遗漏一些东西,但我认为只要“req_text”可以足够独特以仅检测您想要的那些位置就可以了。如果这种方法有任何问题,那么我们当然可以看看替代方法。
猜你喜欢
  • 2019-02-21
  • 2022-01-21
  • 1970-01-01
  • 1970-01-01
  • 2015-03-11
  • 1970-01-01
  • 1970-01-01
  • 2015-06-12
  • 1970-01-01
相关资源
最近更新 更多