【问题标题】:Use object detection of tensorflow to train my model,the time stamp of the ckpt file has not changed for more than 4 hours使用tensorflow的object detection训练我的模型,ckpt文件的时间戳超过4小时没有变化
【发布时间】:2020-04-10 14:11:43
【问题描述】:

我在 Windows 10 上使用 anaconda(python3.6) 和 tensorflow(1.9.0) 来训练我的模型。

我用这个命令训练模型:

python model_main.py --pipeline_config_path=training/ssd_mobilenet_v1_coco.config --model_dir=training/ --num_train_steps=500 --alsologtostderr

Anaconda 提示符输出以下信息。

在我的模型文件夹中,ckpt 文件的时间戳从未改变。

ssd_mobilenet_v1_coco.config 里面的内容是这样的:

# SSD with Mobilenet v1 configuration for MSCOCO Dataset.
# Users should configure the fine_tune_checkpoint field in the train config as
# well as the label_map_path and input_path fields in the train_input_reader and
# eval_input_reader. Search for "PATH_TO_BE_CONFIGURED" to find the fields that
# should be configured.

model {
  ssd {
    num_classes: 2
    box_coder {
      faster_rcnn_box_coder {
        y_scale: 10.0
        x_scale: 10.0
        height_scale: 5.0
        width_scale: 5.0
      }
    }
    matcher {
      argmax_matcher {
        matched_threshold: 0.5
        unmatched_threshold: 0.5
        ignore_thresholds: false
        negatives_lower_than_unmatched: true
        force_match_for_each_row: true
      }
    }
    similarity_calculator {
      iou_similarity {
      }
    }
    anchor_generator {
      ssd_anchor_generator {
        num_layers: 6
        min_scale: 0.2
        max_scale: 0.95
        aspect_ratios: 1.0
        aspect_ratios: 2.0
        aspect_ratios: 0.5
        aspect_ratios: 3.0
        aspect_ratios: 0.3333
      }
    }
    image_resizer {
      fixed_shape_resizer {
        height: 300
        width: 300
      }
    }
    box_predictor {
      convolutional_box_predictor {
        min_depth: 0
        max_depth: 0
        num_layers_before_predictor: 0
        use_dropout: false
        dropout_keep_probability: 0.8
        kernel_size: 1
        box_code_size: 4
        apply_sigmoid_to_scores: false
        conv_hyperparams {
          activation: RELU_6,
          regularizer {
            l2_regularizer {
              weight: 0.00004
            }
          }
          initializer {
            truncated_normal_initializer {
              stddev: 0.03
              mean: 0.0
            }
          }
          batch_norm {
            train: true,
            scale: true,
            center: true,
            decay: 0.9997,
            epsilon: 0.001,
          }
        }
      }
    }
    feature_extractor {
      type: 'ssd_mobilenet_v1'
      min_depth: 16
      depth_multiplier: 1.0
      conv_hyperparams {
        activation: RELU_6,
        regularizer {
          l2_regularizer {
            weight: 0.00004
          }
        }
        initializer {
          truncated_normal_initializer {
            stddev: 0.03
            mean: 0.0
          }
        }
        batch_norm {
          train: true,
          scale: true,
          center: true,
          decay: 0.9997,
          epsilon: 0.001,
        }
      }
    }
    loss {
      classification_loss {
        weighted_sigmoid {
        }
      }
      localization_loss {
        weighted_smooth_l1 {
        }
      }
      hard_example_miner {
        num_hard_examples: 3000
        iou_threshold: 0.99
        loss_type: CLASSIFICATION
        max_negatives_per_positive: 3
        min_negatives_per_image: 0
      }
      classification_weight: 1.0
      localization_weight: 1.0
    }
    normalize_loss_by_num_matches: true
    post_processing {
      batch_non_max_suppression {
        score_threshold: 1e-8
        iou_threshold: 0.6
        max_detections_per_class: 100
        max_total_detections: 100
      }
      score_converter: SIGMOID
    }
  }
}

train_config: {
  batch_size: 1
  optimizer {
    rms_prop_optimizer: {
      learning_rate: {
        exponential_decay_learning_rate {
          initial_learning_rate: 0.004
          decay_steps: 800720
          decay_factor: 0.95
        }
      }
      momentum_optimizer_value: 0.9
      decay: 0.9
      epsilon: 1.0
    }
  }
  #fine_tune_checkpoint: "PATH_TO_BE_CONFIGURED/model.ckpt"
  #from_detection_checkpoint: true
  # Note: The below line limits the training process to 200K steps, which we
  # empirically found to be sufficient enough to train the pets dataset. This
  # effectively bypasses the learning rate schedule (the learning rate will
  # never decay). Remove the below line to train indefinitely.
  num_steps: 100
  data_augmentation_options {
    random_horizontal_flip {
    }
  }
  data_augmentation_options {
    ssd_random_crop {
    }
  }
}

train_input_reader: {
  tf_record_input_reader {
    input_path:'data/train.record'
  }
  label_map_path:'data/side_vehicle.pbtxt'
}

eval_config: {
  num_examples: 8000
  # Note: The below line limits the evaluation process to 10 evaluations.
  # Remove the below line to evaluate indefinitely.
  max_evals: 10
}

eval_input_reader: {
  tf_record_input_reader {
    input_path: 'data/test.record'
  }
  label_map_path: 'data/side_vehicle.pbtxt'
  shuffle: false
  num_readers: 1
}

为什么模型文件的时间戳不变?哪里出错了?

当我使用这个命令训练时:

python model_main.py --pipeline_config_path=training/ssd_mobilenet_v1_coco.config --model_dir=training/ --num_train_steps=10000

错误信息是:

Traceback(最近一次调用最后一次):文件 "E:\Anaconda3\lib\site-packages\tensorflow\python\client\session.py", 第 1334 行,在 _do_call 返回 fn(*args) 文件 "E:\Anaconda3\lib\site-packages\tensorflow\python\client\session.py", 第 1319 行,在 _run_fn 中 选项、feed_dict、fetch_list、target_list、run_metadata)文件“E:\Anaconda3\lib\site-packages\tensorflow\python\client\session.py”, 第 1407 行,在 _call_tf_sessionrun run_metadata) tensorflow.python.framework.errors_impl.InvalidArgumentError: 断言失败:[最大框坐标值大于 1.100000:] [1.11401868] [[{{node ToAbsoluteCoordinates_1/Assert/AssertGuard/Assert}} = Assert[T=[DT_STRING, DT_FLOAT], summarize=3, _device="/job:localhost/replica:0/task:0/device:CPU:0 "](ToAbsoluteCoordinates_1/Assert/AssertGuard/Assert/Switch, ToAbsoluteCoordina tes_1/Assert/AssertGuard/Assert/data_0, ToAbsoluteCoordinates_1/Assert/AssertGuard/Assert/Switch_1)]]

【问题讨论】:

  • 您将步数 num_steps 设置为 100。因此它在第 100 步创建了一个检查点。无论您何时运行(不删除上述位置的检查点文件),它都假定训练已经完成。尝试增加 num_steps。
  • @venkatakrishnan 我已经更改了 num_train_steps=10000,但是错误信息改变了,我已经在我的帖子中更新了。
  • 此错误与您的边界框有关,请检查所有边界框尺寸是否在范围内,即在图像的宽度和高度范围内。它没有更早地抛出这个错误,因为在前 100 步中它没有通过有错误的框。
  • @venkatakrishnan 你的意思是我应该检查 train.csv?在我的 train.csv 中,内容如下: 文件名 width height class xmin ymin xmax ymax 1.jpg 1138 677 side 127 560 340 594 2.jpg 1138 677 side 339 561 549 595 3.jpg 1138 677 side 548 562 759 597 4 .jpg 1138 677 侧 758 564 932 596 5.jpg 1166 697 侧 75 528 273 565 6.jpg 1161 646 侧 451 544 620 573
  • 是的。基本上你的 xml 文件带有边界框。寻找一些提供 python 脚本来检查边界框的 github 存储库。

标签: python tensorflow computer-vision anaconda object-detection


【解决方案1】:

现在我知道哪里错了。我的 tensorflow 版本是 1.9.0。我把tensorflow的版本改成1.12.0,然后修改了这个文件box_list_ops.py,设置check_range=False,问题就解决了。

【讨论】:

    猜你喜欢
    • 2017-12-20
    • 1970-01-01
    • 2018-07-28
    • 1970-01-01
    • 2019-07-18
    • 2018-01-14
    • 2018-06-17
    • 2019-06-17
    • 2020-10-16
    相关资源
    最近更新 更多